Words That Work, Tested: the word-pair counts
Every term the desk counted for the Book Club piece, by the desk’s own outlet-lean bucket, with denominators, the quotation-mark split, the script that produced them and the hash of that script and its output.
Universe
- 1013 live runs (approved.json minus retired.json); 729 are coverage or delta runs; 723 have a corpus; first 2026-06-17T08-52-43Z, last 2026-10-07T05-53-48Z.
- 8,190 distinct articles counted, de-duplicated by URL. Skipped: {"dup_url": 982, "nourl_or_nobody": 168, "desk": 1}.
- Editorials, audits and dispatches freeze primary documents and are left out; model-written and operator sources are left out.
Buckets and denominators
The desk’s map (parrot.sitegen _OUTLET_LEAN, AllSides ratings): Fox News = R, Townhall = R, Fox Business = LR, CBS News = LL, Associated Press = LL, NPR = LL, CNN = LL, NBC News = LL, Politico = LL, Time = LL, The Guardian = LL, Mediaite = LL, The Boston Globe = LL, The Philadelphia Inquirer = LL, MS NOW = L, Democracy Now = L, Newsweek = C, The Hill = C, BBC News = C, Reuters = C, Axios = C. Largest outlets per bucket:
- Left: MS NOW 68, Democracy Now 2
- Lean Left: Associated Press 377, CNN 210, CBS News 185, NBC News 144, Politico 124, NPR 113
- Center: The Hill 400, Reuters 292, Newsweek 196, BBC News 170, Axios 79
- Lean Right: Fox Business 39
- Right: Fox News 331, Townhall 54
- Unrated: Washington Examiner 299, Al Jazeera 234, New York Post 198, The Independent (UK) 195, Breitbart 172, USA Today 160
Counts by pair
death tax vs estate tax
Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| death tax | 1 (0) | 0 (0) | 1 (0) | 0 (0) | 0 (0) | 0 (0) | 2 (0) |
| estate tax | 0 (0) | 1 (1) | 0 (0) | 0 (0) | 0 (0) | 3 (2) | 4 (3) |
| inheritance tax | 0 (0) | 1 (1) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 1 (1) |
climate change vs global warming
Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| climate change | 1 (0) | 31 (30) | 23 (20) | 0 (0) | 4 (3) | 71 (61) | 130 (114) |
| global warming | 0 (0) | 2 (1) | 3 (2) | 0 (0) | 0 (0) | 17 (12) | 22 (15) |
| climate crisis/emergency | 0 (0) | 2 (1) | 1 (1) | 0 (0) | 0 (0) | 13 (7) | 16 (9) |
drilling vs exploration (oil and gas)
Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| oil/gas drilling | 0 (0) | 2 (2) | 1 (1) | 0 (0) | 0 (0) | 8 (8) | 11 (11) |
| oil/gas exploration | 0 (0) | 0 (0) | 1 (1) | 0 (0) | 0 (0) | 3 (3) | 4 (4) |
government takeover vs government-run
Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| government takeover | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 1 (1) | 1 (1) | 2 (2) |
| government-run | 0 (0) | 2 (2) | 2 (2) | 0 (0) | 0 (0) | 7 (7) | 11 (11) |
illegal alien / illegal immigrant / undocumented
Articles using the term at least once, with the count outside quotation marks in brackets. Vocabulary pair only; the desk did not source it to Luntz’s own words in this run.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| illegal alien(s) | 0 (0) | 11 (2) | 8 (0) | 1 (0) | 7 (1) | 45 (12) | 72 (15) |
| illegal immigrant(s) | 0 (0) | 4 (3) | 5 (1) | 0 (0) | 7 (7) | 21 (18) | 37 (29) |
| undocumented | 0 (0) | 11 (10) | 7 (7) | 0 (0) | 2 (2) | 31 (28) | 51 (47) |
tax relief vs tax cut
Articles using the term at least once, with the count outside quotation marks in brackets. Vocabulary pair only; the desk did not source it to Luntz’s own words in this run.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| tax relief | 0 (0) | 3 (3) | 1 (0) | 0 (0) | 1 (1) | 13 (10) | 18 (14) |
| tax cut(s) | 2 (1) | 27 (22) | 9 (6) | 1 (1) | 2 (2) | 79 (61) | 120 (93) |
school voucher vs opportunity scholarship
Articles using the term at least once, with the count outside quotation marks in brackets. Vocabulary pair only; the desk did not source it to Luntz’s own words in this run.
| Term | Left n=70 | Lean Left n=1,261 | Center n=1,137 | Lean Right n=39 | Right n=385 | Unrated n=5,298 | All n=8,190 |
|---|---|---|---|---|---|---|---|
| school voucher(s) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 6 (4) | 6 (4) |
| opportunity scholarship(s) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 0 (0) |
Receipt
sha256 of the script: 7b8abd63d2c8402b8e9f7fc92ca3d5f233c80938b642855bd54b86cbec7c67fb. sha256 of the full operator source (summary lines, tables and script): 072dad375441fee3551acf799f6e544bf05e093f8fc5b1bc053d60de9f337902. Run on the DGX on 2026-10-07 against data/runs/*/corpus.jsonl; deterministic regex counts, no model read an article.
Script
#!/usr/bin/env python3
"""Paired-term usage in the desk's frozen corpora, by outlet lean bucket.
Universe : runs listed in data/operator/approved.json minus data/operator/retired.json whose audit.json kind is
'coverage' or 'delta' (news-story runs; editorials, audits and dispatches freeze primary documents and
are left out) and that have a corpus.jsonl; every source whose kind is not llm/operator/document and
whose outlet is not the desk itself; de-duplicated by URL (one article counted once however many runs
froze it).
Buckets : the desk's own map, parrot.sitegen._OUTLET_LEAN (AllSides ratings; L, LL, C, LR, R), after the desk's
own outlet-name normalisation (parrot.sitegen._norm_outlet). Outlets AllSides does not rate here are
'unrated' and are reported, not dropped.
Counting : an article 'uses' a term if the term occurs at least once in the source body (case-insensitive).
'outside quotes' = the same test after deleting text inside paired quotation marks (curly or straight),
a heuristic, not a parse. Mentions (every occurrence) are also totalled.
Caveat : the corpus is the set of stories the desk chose to cover and the articles it froze for them, not a
random sample of coverage. Counts are small. Outlet style guides differ (AP, for one).
"""
import json, os, re, sys, hashlib, collections
ROOT = os.path.expanduser("~/stochastic-parrot")
sys.path.insert(0, ROOT)
from parrot.sitegen import _norm_outlet, _OUTLET_LEAN, _LEAN_META
ap = set(json.load(open(f"{ROOT}/data/operator/approved.json"))["run_ids"])
rt = set(json.load(open(f"{ROOT}/data/operator/retired.json"))["run_ids"])
live = sorted(ap - rt)
def _kind(r):
try:
return json.load(open(f"{ROOT}/data/runs/{r}/audit.json")).get("kind")
except Exception:
return None
live_all = live
live = [r for r in live_all if _kind(r) in ("coverage", "delta")]
runs = [r for r in live if os.path.exists(f"{ROOT}/data/runs/{r}/corpus.jsonl")]
NONNEWS = {"llm", "operator", "document"}
BUCKETS = ["L", "LL", "C", "LR", "R", "unrated"]
BLABEL = {"L": "Left", "LL": "Lean Left", "C": "Center", "LR": "Lean Right", "R": "Right", "unrated": "Unrated by the desk's map"}
TERMS = collections.OrderedDict([
("death tax", r"\bdeath tax(?:es)?\b"),
("estate tax", r"(?<!real )\bestate tax(?:es)?\b"),
("inheritance tax", r"\binheritance tax(?:es)?\b"),
("climate change", r"\bclimate change\b"),
("global warming", r"\bglobal warming\b"),
("climate crisis/emergency", r"\bclimate (?:crisis|emergency)\b"),
("illegal alien(s)", r"\billegal aliens?\b"),
("illegal immigrant(s)", r"\billegal immigrants?\b"),
("undocumented", r"\bundocumented\b"),
("tax relief", r"\btax relief\b"),
("tax cut(s)", r"\btax cuts?\b"),
("oil/gas drilling", r"\b(?:oil|gas|offshore|arctic|energy) drilling\b|\bdrill(?:ing)? for (?:oil|gas)\b"),
("oil/gas exploration", r"\b(?:oil|gas|offshore|arctic|energy|natural gas|oil and gas) exploration\b|\bexplor(?:e|ing) for (?:oil|gas)\b"),
("government takeover", r"\bgovernment takeover\b"),
("government-run", r"\bgovernment[- ]run\b"),
("school voucher(s)", r"\bschool vouchers?\b|\bvoucher (?:programs?|systems?)\b|\beducation vouchers?\b"),
("opportunity scholarship(s)", r"\bopportunity scholarships?\b"),
])
PAIRS = [
("death tax vs estate tax", ["death tax", "estate tax", "inheritance tax"], True),
("climate change vs global warming", ["climate change", "global warming", "climate crisis/emergency"], True),
("drilling vs exploration (oil and gas)", ["oil/gas drilling", "oil/gas exploration"], True),
("government takeover vs government-run", ["government takeover", "government-run"], True),
("illegal alien / illegal immigrant / undocumented", ["illegal alien(s)", "illegal immigrant(s)", "undocumented"], False),
("tax relief vs tax cut", ["tax relief", "tax cut(s)"], False),
("school voucher vs opportunity scholarship", ["school voucher(s)", "opportunity scholarship(s)"], False),
]
RX = {k: re.compile(v, re.I) for k, v in TERMS.items()}
QRX = [re.compile(r"“[^”]{0,1500}”"), re.compile(r'"[^"\n]{0,1500}"')]
def strip_quotes(t):
for q in QRX:
t = q.sub(" ", t)
return t
def bucket(outlet):
return _OUTLET_LEAN.get(_norm_outlet(outlet), "unrated")
seen = {}
skipped = collections.Counter()
for r in runs:
for line in open(f"{ROOT}/data/runs/{r}/corpus.jsonl", encoding="utf-8"):
line = line.strip()
if not line:
continue
d = json.loads(line)
url = d.get("url") or ""
outlet = d.get("outlet") or ""
kind = (d.get("kind") or "newsroom").lower()
if kind in NONNEWS:
skipped["kind:" + kind] += 1; continue
if "stochastic parrot" in outlet.lower() or outlet.lower().startswith("the desk"):
skipped["desk"] += 1; continue
if not url or not d.get("body"):
skipped["nourl_or_nobody"] += 1; continue
if url in seen:
skipped["dup_url"] += 1; continue
seen[url] = (r, outlet, d["body"])
den = collections.Counter(); out_den = collections.defaultdict(collections.Counter)
use = {k: collections.Counter() for k in TERMS} # articles using term, by bucket
use_nq = {k: collections.Counter() for k in TERMS} # ... outside quotation marks
ment = {k: collections.Counter() for k in TERMS} # total mentions
outlets_by_bucket = collections.defaultdict(collections.Counter)
for url, (run, outlet, body) in seen.items():
b = bucket(outlet); den[b] += 1; den["ALL"] += 1
outlets_by_bucket[b][_norm_outlet(outlet)] += 1
nq = strip_quotes(body)
for k, rx in RX.items():
n = len(rx.findall(body))
if n:
use[k][b] += 1; use[k]["ALL"] += 1; ment[k][b] += n; ment[k]["ALL"] += n
if rx.search(nq):
use_nq[k][b] += 1; use_nq[k]["ALL"] += 1
res = {
"universe": {"live_runs_approved_minus_retired": len(live_all), "of_which_coverage_or_delta": len(live), "with_corpus": len(runs),
"articles_counted_deduped_by_url": den["ALL"], "skipped": dict(skipped),
"run_id_first": runs[0], "run_id_last": runs[-1]},
"bucket_map": {k: v for k, v in _OUTLET_LEAN.items()},
"denominators": {b: den[b] for b in BUCKETS + ["ALL"]},
"top_outlets_by_bucket": {b: outlets_by_bucket[b].most_common(6) for b in BUCKETS},
"terms": {k: {"articles": {b: use[k][b] for b in BUCKETS + ["ALL"]},
"articles_outside_quotes": {b: use_nq[k][b] for b in BUCKETS + ["ALL"]},
"mentions": {b: ment[k][b] for b in BUCKETS + ["ALL"]}} for k in TERMS},
"pairs": [{"name": n, "terms": t, "luntz_sourced": ls} for n, t, ls in PAIRS],
}
json.dump(res, open(sys.argv[1] if len(sys.argv) > 1 else "/dev/stdout", "w"), indent=1, ensure_ascii=False)
def pct(a, b): return f"{100*a/b:.1f}%" if b else "n/a"
lines = []
lines.append(f"Universe: {len(live_all)} runs in approved.json minus retired.json; {len(live)} are coverage or delta runs; {len(runs)} have a corpus; "
f"{den['ALL']} distinct articles counted (de-duplicated by URL); skipped: {dict(skipped)}")
lines.append("Denominators (articles): " + ", ".join(f"{BLABEL[b]} {den[b]}" for b in BUCKETS) + f", all {den['ALL']}")
for name, terms, ls in PAIRS:
lines.append("")
lines.append(f"== {name} (articles using the term / articles in bucket; outside-quotes in brackets)")
hdr = "term".ljust(28) + "".join(b.rjust(14) for b in BUCKETS) + "ALL".rjust(14)
lines.append(hdr)
for t in terms:
row = t.ljust(28)
for b in BUCKETS + ["ALL"]:
row += f"{use[t][b]} [{use_nq[t][b]}]/{den[b]}".rjust(14)
lines.append(row)
txt = "\n".join(lines)
print(txt)
open((sys.argv[1] if len(sys.argv) > 1 else "/tmp/out.json") + ".txt", "w").write(txt)