← Words That Work, Tested

Words That Work, Tested: the word-pair counts

Every term the desk counted for the Book Club piece, by the desk’s own outlet-lean bucket, with denominators, the quotation-mark split, the script that produced them and the hash of that script and its output.

Scope and caveats: a convenience sample of the stories the desk covered, not a random sample of coverage. Counts are small. The desk’s outlet map tags only outlets that AllSides rates and only some of those, so 5,298 of 8,190 articles are Unrated; they are counted, not dropped. Outlet style guides (the Associated Press’s among them) settle some word choices before a reporter does. “Outside quotation marks” is a heuristic (text between paired quotation marks is removed), not a parse. These counts show vocabulary in these articles. They do not show persuasion, and nothing here says Luntz caused any usage.

Universe

Buckets and denominators

The desk’s map (parrot.sitegen _OUTLET_LEAN, AllSides ratings): Fox News = R, Townhall = R, Fox Business = LR, CBS News = LL, Associated Press = LL, NPR = LL, CNN = LL, NBC News = LL, Politico = LL, Time = LL, The Guardian = LL, Mediaite = LL, The Boston Globe = LL, The Philadelphia Inquirer = LL, MS NOW = L, Democracy Now = L, Newsweek = C, The Hill = C, BBC News = C, Reuters = C, Axios = C. Largest outlets per bucket:

Counts by pair

death tax vs estate tax

Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
death tax1 (0)0 (0)1 (0)0 (0)0 (0)0 (0)2 (0)
estate tax0 (0)1 (1)0 (0)0 (0)0 (0)3 (2)4 (3)
inheritance tax0 (0)1 (1)0 (0)0 (0)0 (0)0 (0)1 (1)

climate change vs global warming

Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
climate change1 (0)31 (30)23 (20)0 (0)4 (3)71 (61)130 (114)
global warming0 (0)2 (1)3 (2)0 (0)0 (0)17 (12)22 (15)
climate crisis/emergency0 (0)2 (1)1 (1)0 (0)0 (0)13 (7)16 (9)

drilling vs exploration (oil and gas)

Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
oil/gas drilling0 (0)2 (2)1 (1)0 (0)0 (0)8 (8)11 (11)
oil/gas exploration0 (0)0 (0)1 (1)0 (0)0 (0)3 (3)4 (4)

government takeover vs government-run

Articles using the term at least once, with the count outside quotation marks in brackets. Tied to a Luntz-sourced claim in the piece.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
government takeover0 (0)0 (0)0 (0)0 (0)1 (1)1 (1)2 (2)
government-run0 (0)2 (2)2 (2)0 (0)0 (0)7 (7)11 (11)

illegal alien / illegal immigrant / undocumented

Articles using the term at least once, with the count outside quotation marks in brackets. Vocabulary pair only; the desk did not source it to Luntz’s own words in this run.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
illegal alien(s)0 (0)11 (2)8 (0)1 (0)7 (1)45 (12)72 (15)
illegal immigrant(s)0 (0)4 (3)5 (1)0 (0)7 (7)21 (18)37 (29)
undocumented0 (0)11 (10)7 (7)0 (0)2 (2)31 (28)51 (47)

tax relief vs tax cut

Articles using the term at least once, with the count outside quotation marks in brackets. Vocabulary pair only; the desk did not source it to Luntz’s own words in this run.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
tax relief0 (0)3 (3)1 (0)0 (0)1 (1)13 (10)18 (14)
tax cut(s)2 (1)27 (22)9 (6)1 (1)2 (2)79 (61)120 (93)

school voucher vs opportunity scholarship

Articles using the term at least once, with the count outside quotation marks in brackets. Vocabulary pair only; the desk did not source it to Luntz’s own words in this run.

TermLeft
n=70
Lean Left
n=1,261
Center
n=1,137
Lean Right
n=39
Right
n=385
Unrated
n=5,298
All
n=8,190
school voucher(s)0 (0)0 (0)0 (0)0 (0)0 (0)6 (4)6 (4)
opportunity scholarship(s)0 (0)0 (0)0 (0)0 (0)0 (0)0 (0)0 (0)

Receipt

sha256 of the script: 7b8abd63d2c8402b8e9f7fc92ca3d5f233c80938b642855bd54b86cbec7c67fb. sha256 of the full operator source (summary lines, tables and script): 072dad375441fee3551acf799f6e544bf05e093f8fc5b1bc053d60de9f337902. Run on the DGX on 2026-10-07 against data/runs/*/corpus.jsonl; deterministic regex counts, no model read an article.

Script

#!/usr/bin/env python3
"""Paired-term usage in the desk's frozen corpora, by outlet lean bucket.

Universe : runs listed in data/operator/approved.json minus data/operator/retired.json whose audit.json kind is
           'coverage' or 'delta' (news-story runs; editorials, audits and dispatches freeze primary documents and
           are left out) and that have a corpus.jsonl; every source whose kind is not llm/operator/document and
           whose outlet is not the desk itself; de-duplicated by URL (one article counted once however many runs
           froze it).
Buckets  : the desk's own map, parrot.sitegen._OUTLET_LEAN (AllSides ratings; L, LL, C, LR, R), after the desk's
           own outlet-name normalisation (parrot.sitegen._norm_outlet). Outlets AllSides does not rate here are
           'unrated' and are reported, not dropped.
Counting : an article 'uses' a term if the term occurs at least once in the source body (case-insensitive).
           'outside quotes' = the same test after deleting text inside paired quotation marks (curly or straight),
           a heuristic, not a parse. Mentions (every occurrence) are also totalled.
Caveat   : the corpus is the set of stories the desk chose to cover and the articles it froze for them, not a
           random sample of coverage. Counts are small. Outlet style guides differ (AP, for one).
"""
import json, os, re, sys, hashlib, collections
ROOT = os.path.expanduser("~/stochastic-parrot")
sys.path.insert(0, ROOT)
from parrot.sitegen import _norm_outlet, _OUTLET_LEAN, _LEAN_META

ap = set(json.load(open(f"{ROOT}/data/operator/approved.json"))["run_ids"])
rt = set(json.load(open(f"{ROOT}/data/operator/retired.json"))["run_ids"])
live = sorted(ap - rt)
def _kind(r):
    try:
        return json.load(open(f"{ROOT}/data/runs/{r}/audit.json")).get("kind")
    except Exception:
        return None
live_all = live
live = [r for r in live_all if _kind(r) in ("coverage", "delta")]
runs = [r for r in live if os.path.exists(f"{ROOT}/data/runs/{r}/corpus.jsonl")]

NONNEWS = {"llm", "operator", "document"}
BUCKETS = ["L", "LL", "C", "LR", "R", "unrated"]
BLABEL = {"L": "Left", "LL": "Lean Left", "C": "Center", "LR": "Lean Right", "R": "Right", "unrated": "Unrated by the desk's map"}

TERMS = collections.OrderedDict([
    ("death tax",            r"\bdeath tax(?:es)?\b"),
    ("estate tax",           r"(?<!real )\bestate tax(?:es)?\b"),
    ("inheritance tax",      r"\binheritance tax(?:es)?\b"),
    ("climate change",       r"\bclimate change\b"),
    ("global warming",       r"\bglobal warming\b"),
    ("climate crisis/emergency", r"\bclimate (?:crisis|emergency)\b"),
    ("illegal alien(s)",     r"\billegal aliens?\b"),
    ("illegal immigrant(s)", r"\billegal immigrants?\b"),
    ("undocumented",         r"\bundocumented\b"),
    ("tax relief",           r"\btax relief\b"),
    ("tax cut(s)",           r"\btax cuts?\b"),
    ("oil/gas drilling",     r"\b(?:oil|gas|offshore|arctic|energy) drilling\b|\bdrill(?:ing)? for (?:oil|gas)\b"),
    ("oil/gas exploration",  r"\b(?:oil|gas|offshore|arctic|energy|natural gas|oil and gas) exploration\b|\bexplor(?:e|ing) for (?:oil|gas)\b"),
    ("government takeover",  r"\bgovernment takeover\b"),
    ("government-run",       r"\bgovernment[- ]run\b"),
    ("school voucher(s)",    r"\bschool vouchers?\b|\bvoucher (?:programs?|systems?)\b|\beducation vouchers?\b"),
    ("opportunity scholarship(s)", r"\bopportunity scholarships?\b"),
])
PAIRS = [
    ("death tax vs estate tax", ["death tax", "estate tax", "inheritance tax"], True),
    ("climate change vs global warming", ["climate change", "global warming", "climate crisis/emergency"], True),
    ("drilling vs exploration (oil and gas)", ["oil/gas drilling", "oil/gas exploration"], True),
    ("government takeover vs government-run", ["government takeover", "government-run"], True),
    ("illegal alien / illegal immigrant / undocumented", ["illegal alien(s)", "illegal immigrant(s)", "undocumented"], False),
    ("tax relief vs tax cut", ["tax relief", "tax cut(s)"], False),
    ("school voucher vs opportunity scholarship", ["school voucher(s)", "opportunity scholarship(s)"], False),
]
RX = {k: re.compile(v, re.I) for k, v in TERMS.items()}
QRX = [re.compile(r"“[^”]{0,1500}”"), re.compile(r'"[^"\n]{0,1500}"')]

def strip_quotes(t):
    for q in QRX:
        t = q.sub(" ", t)
    return t

def bucket(outlet):
    return _OUTLET_LEAN.get(_norm_outlet(outlet), "unrated")

seen = {}
skipped = collections.Counter()
for r in runs:
    for line in open(f"{ROOT}/data/runs/{r}/corpus.jsonl", encoding="utf-8"):
        line = line.strip()
        if not line:
            continue
        d = json.loads(line)
        url = d.get("url") or ""
        outlet = d.get("outlet") or ""
        kind = (d.get("kind") or "newsroom").lower()
        if kind in NONNEWS:
            skipped["kind:" + kind] += 1; continue
        if "stochastic parrot" in outlet.lower() or outlet.lower().startswith("the desk"):
            skipped["desk"] += 1; continue
        if not url or not d.get("body"):
            skipped["nourl_or_nobody"] += 1; continue
        if url in seen:
            skipped["dup_url"] += 1; continue
        seen[url] = (r, outlet, d["body"])

den = collections.Counter(); out_den = collections.defaultdict(collections.Counter)
use = {k: collections.Counter() for k in TERMS}      # articles using term, by bucket
use_nq = {k: collections.Counter() for k in TERMS}   # ... outside quotation marks
ment = {k: collections.Counter() for k in TERMS}     # total mentions
outlets_by_bucket = collections.defaultdict(collections.Counter)
for url, (run, outlet, body) in seen.items():
    b = bucket(outlet); den[b] += 1; den["ALL"] += 1
    outlets_by_bucket[b][_norm_outlet(outlet)] += 1
    nq = strip_quotes(body)
    for k, rx in RX.items():
        n = len(rx.findall(body))
        if n:
            use[k][b] += 1; use[k]["ALL"] += 1; ment[k][b] += n; ment[k]["ALL"] += n
            if rx.search(nq):
                use_nq[k][b] += 1; use_nq[k]["ALL"] += 1

res = {
    "universe": {"live_runs_approved_minus_retired": len(live_all), "of_which_coverage_or_delta": len(live), "with_corpus": len(runs),
                 "articles_counted_deduped_by_url": den["ALL"], "skipped": dict(skipped),
                 "run_id_first": runs[0], "run_id_last": runs[-1]},
    "bucket_map": {k: v for k, v in _OUTLET_LEAN.items()},
    "denominators": {b: den[b] for b in BUCKETS + ["ALL"]},
    "top_outlets_by_bucket": {b: outlets_by_bucket[b].most_common(6) for b in BUCKETS},
    "terms": {k: {"articles": {b: use[k][b] for b in BUCKETS + ["ALL"]},
                  "articles_outside_quotes": {b: use_nq[k][b] for b in BUCKETS + ["ALL"]},
                  "mentions": {b: ment[k][b] for b in BUCKETS + ["ALL"]}} for k in TERMS},
    "pairs": [{"name": n, "terms": t, "luntz_sourced": ls} for n, t, ls in PAIRS],
}
json.dump(res, open(sys.argv[1] if len(sys.argv) > 1 else "/dev/stdout", "w"), indent=1, ensure_ascii=False)

def pct(a, b): return f"{100*a/b:.1f}%" if b else "n/a"
lines = []
lines.append(f"Universe: {len(live_all)} runs in approved.json minus retired.json; {len(live)} are coverage or delta runs; {len(runs)} have a corpus; "
             f"{den['ALL']} distinct articles counted (de-duplicated by URL); skipped: {dict(skipped)}")
lines.append("Denominators (articles): " + ", ".join(f"{BLABEL[b]} {den[b]}" for b in BUCKETS) + f", all {den['ALL']}")
for name, terms, ls in PAIRS:
    lines.append("")
    lines.append(f"== {name}  (articles using the term / articles in bucket; outside-quotes in brackets)")
    hdr = "term".ljust(28) + "".join(b.rjust(14) for b in BUCKETS) + "ALL".rjust(14)
    lines.append(hdr)
    for t in terms:
        row = t.ljust(28)
        for b in BUCKETS + ["ALL"]:
            row += f"{use[t][b]} [{use_nq[t][b]}]/{den[b]}".rjust(14)
        lines.append(row)
txt = "\n".join(lines)
print(txt)
open((sys.argv[1] if len(sys.argv) > 1 else "/tmp/out.json") + ".txt", "w").write(txt)