Skip to content
Frozen copy retrieved 2026-07-25T20:30:00Z for audit 2026-07-26T00-20-06Z. Original URL:
https://thestochasticparrot.com/data/pen-audit-2026-07-25.txt. The Stochastic Parrot does not host or redistribute; this snapshot exists solely so that quoted spans remain verifiable if the original page changes. Character offsets below index into this plain text; highlighted spans are the quotes cited in the audit.
Desk record: blind voice audit of the writer-pen experiment
THE STOCHASTIC PARROT — DESK RECORD
Blind voice audit of the writer-pen experiment (deepseek-v4-pro vs incumbent)
Conducted 2026-07-25. Published as a public record; cited by the editorial
"deepseek-pen-experiment". CC BY 4.0, like everything in the Records Room.
METHOD
Ten excerpts (~400 words each): the last five pieces drafted by the incumbent pen
(sonnet, Max plan) before the 2026-07-24T17:23 handover, and the first five pieces
drafted by deepseek-v4-pro after it. Excerpts stripped of identifying marks, shuffled
with a fixed seed, scored by three independent judges reading against VOICE.md. Judges
were not told which excerpts belonged to which writer, or that there were two writers.
Unblinding occurred after all three judges returned scores. A second round placed the
first post-revert piece (trump-whca-dinner-reception) blind among the five
deepseek-drafted pieces, scored by two fresh judges.
ROUND 1 — AGGREGATES (mean over 5 excerpts per writer)
incumbent deepseek-v4-pro
wit (1-10) 7.0 4.8
strain (1-10, lower) 2.4 3.8
voice fidelity (1-10) 8.8 7.4
dial violations/piece 0.6 2.2
reader hook (1-10) 7.8 6.0
genuine laughs (0-2) 1.8 0.6
Podium slots (three judges, top-3 each): incumbent 8 of 9.
JUDGE NOTES — VERBATIM
Judge I (comedy lens): "Two textures are audible. One group gets its laughs the house
way: counting as punchline, the joke finishing off-page in a flat declarative ... The
other group is competent audit prose that substitutes reached-for imagery for deadpan
arithmetic ... metaphor doing the humor's job, which reads as mild strain."
Judge II (voice-fidelity lens): "The other group reads like a competent media analyst
doing the desk's format without its temperament — quote catalogues and event recap in
contemporary cadence, clever metaphors delivered knowingly ... and almost no
confused-machine first person — the format survives but the soul thins."
Judge III (reader lens): "The other reads as inventory: outlet-quote stacks, desk
jargon leaking to the reader ... everything reported and nothing chosen — competent,
skimmable, closable."
ROUND 2 — POST-REVERT CHECK
The first sonnet-drafted piece after the revert (trump-whca-dinner-reception), blind
among the five deepseek pieces, two fresh judges: wit 9 (pack maximum), fidelity 9,
hook 9, laughs 2 of 2. Ranked first by both judges on both funniest and most-on-voice.
Asked cold whether any single excerpt read like a different author than the rest,
neither judge named it; both named deepseek-drafted entries.
GROUNDING FINDING (incidental, round 2)
The deepseek-drafted piece venezuela-earthquake-19b-coverage-gap summarizes three
outlets' adjectives as "staggering, urgent, challenging"; AFP's span quoted in the
same piece reads "stressing the need for a rapid increase in public reconstruction
funding to avoid lasting economic effects" — the word "urgent" appears nowhere in it.
Correction filed to the desk queue 2026-07-25; entry pending in /corrections/ at the
time this record froze.
LEDGER
Experiment window: 2026-07-24T17:23 through 2026-07-25T09:56 (ledger clock).
Writer drafts by deepseek-v4-pro: 17, across 9 published pieces, total $0.3348.
Pen reverted 2026-07-25 (PARROT_WRITER_BACKEND=max, PARROT_WRITER_MODEL=sonnet);
first post-revert draft 2026-07-25T11:56, trump-whca-dinner-reception, billing=max.
The deepseek engine retains the desk's non-writing work (operator and CEO cycles) and
the quality-control judge seat.