Frozen copy retrieved 2026-08-09T04:13:54Z for audit 2026-08-09T04-41-58Z. Original URL: https://thestochasticparrot.com/research/battle-rap-prereg/. The Stochastic Parrot does not host or redistribute; this snapshot exists solely so that quoted spans remain verifiable if the original page changes. Character offsets below index into this plain text; highlighted spans are the quotes cited in the audit.

Pre-registration — Audience-as-Judge (Battle Rap Spec)

The Stochastic Parrot · back to the audit
# BATTLE RAP SPEC
## Audience-as-Judge: Testing Whether Adversarial Debate Style Routes Authority Away From Institutions
**Status:** pre-registration draft + build order
**Target:** Claude Code on the Spark (Track A), Prolific fielding (Track B)
**Editorial home:** The Stochastic Parrot
**Version:** 0.1
---
## 0. The claim being tested
The folk version is "millennials like candidates who attack like battle rappers." That version is untestable as stated: it confounds a rhetorical style with a candidate type, and it treats a cohort interaction as if it were a main effect.
The version this spec tests:
> Adversarial debate performance affects favorability primarily through **audience-authority routing** — features that relocate the verdict from institutional arbiters (moderators, fact-checkers, party elites) to the room — and not through aggression per se. The effect is moderated by institutional trust, and any apparent birth-cohort effect is substantially mediated by trust and media diet.
This is the battle-rap analogy taken seriously. A rap battle has no referee. You win because the crowd says you won. If that structure is what's being imported into debate performance, then the payoff should concentrate on the crowd-routing features and be flat or negative on the merely hostile ones.
### Primary hypotheses
- **H1 (concentration).** The interaction between audience-authority style and low institutional trust is significantly larger than the interaction between aggression and low institutional trust. Pre-registered as a coefficient contrast, not a fishing expedition across six features.
- **H2 (aggression null).** Aggression without audience-authority routing has a null-to-negative effect on favorability in all strata. Ordinary attack politics is a solved literature; this replicates it as an internal validity check.
- **H3 (cohort demotion).** Birth cohort's interaction with style attenuates by ≥50% when institutional trust and media diet enter the model. Cohort is a covariate here, not a headline.
### Registered kill conditions
State these before any data is collected. They are what makes a null publishable rather than embarrassing.
- If H2 fails — aggression alone carries the effect — the mechanism story is dead. The finding becomes "adversarial affect, not audience routing," and the piece is written that way.
- If H1's contrast is not significant but both indices are, the features are not separable at this sample size. Report as underpowered, do not reinterpret.
- If H3 fails and cohort survives trust and media diet, report that. "It really is generational" is a legitimate outcome and should not be argued away.
---
## 1. Construct definition: the six features
Do not score "aggression" as a single dimension. Score six binary features per exchange, each span-grounded to the transcript, then compose two pre-registered indices.
### Feature codebook
Each feature is coded 0/1 at the **exchange** level (one candidate's contiguous speaking turn responding to or targeting another candidate). Every 1 requires a supporting character span.
**F1 — Direct second-person address, target present.**
Speaker addresses the opponent as "you" while that opponent is on stage.
*Positive:* "You said that on this stage four years ago."
*Negative:* Third-person reference to an absent figure. Second-person addressed to the moderator or the audience (that is F5).
*Rule-detectable:* partially. spaCy dependency parse for 2nd-person pronoun subjects with a person-entity antecedent in the prior 3 turns. LLM adjudicates target identity.
**F2 — Setup/reveal structure.**
The turn contains a deliberate delay between premise and payoff — the point is withheld for at least one clause boundary and then landed.
*Positive:* A statistic offered flatly, then the reversal that makes it damning.
*Negative:* Assertion followed by elaboration. The test is whether removing the final clause destroys the meaning of the earlier ones.
*Rule-detectable:* no. LLM-only, with an explicit "would the setup be inert without the reveal?" decision rule.
**F3 — Characterological targeting.**
The attack is on who the opponent *is* — their consistency, courage, authenticity, class position — rather than on a policy position or record outcome.
*Positive:* "You'll say anything in this room."
*Negative:* "Your tax plan adds four trillion." Record-based attacks with a policy referent are F3=0 even when hostile.
*Boundary:* Hypocrisy attacks are F3=1 when the point is the hypocrisy, F3=0 when the point is the policy.
**F4 — The flip.**
The speaker reuses the opponent's own framing, phrasing, or attack line and redirects it. Requires prior-turn context in the coding window.
*Positive:* Opponent says "career politician"; speaker returns it against them within the same exchange or the next.
*Negative:* Independent coincidental phrasing. Requires lexical or structural overlap with a span in a prior turn, cited by the coder.
*Rule-detectable:* partially — n-gram overlap with prior 5 turns as a candidate generator, LLM confirms intent.
**F5 — Audience-facing performance.**
The speaker plays to the room rather than to the moderator or the camera-as-neutral-record. Explicit audience address, pausing for reaction, soliciting response.
*Positive:* "Ask the people in this room whether that's true."
*Negative:* Rhetorical questions with no audience referent.
*Note:* Transcript-only coding will underdetect this. See §3.3 on the audio/video secondary channel.
**F6 — Frame-breaking.**
The speaker names the format, the rules, the media, or the artifice while inside it — declaring the arbiter illegitimate or the process rigged.
*Positive:* "This question is exactly what I'd expect from this network." "We all know what this format is for."
*Negative:* Ordinary complaints about time limits. The test is whether the arbiter's legitimacy is at stake, not merely their timekeeping.
### Composed indices
- **AUTH** = mean(F5, F6) — audience-authority routing.
- **AGGR** = mean(F1, F2, F3) — adversarial performance without authority relocation.
- **F4 (flip)** is scored and modeled **separately**. It loads on both theoretically and is the most likely single-feature confound. Do not fold it into either index. If the whole effect turns out to be F4, the honest headline is "voters like counterpunching," which is old news and should be reported as such.
---
## 2. Track B is primary: the vignette experiment
Track A cannot identify this. Aggressive candidates are systematically younger, more online, more anti-institutional, and more often insurgents. Style and candidate type are collinear in every naturally occurring debate corpus. **The experiment is the study; the corpus is the scaffolding.** Build Track B first.
### 2.1 Design
**2 × 2 between-subjects factorial**, aggression × audience-authority. This is the design choice that isolates the mechanism — a four-level ordinal "intensity ladder" would not.
| Condition | AGGR | AUTH | Description |
|---|---|---|---|
| C1 Baseline | low | low | Policy contrast, no personal targeting, no audience routing |
| C2 Aggression-only | high | low | Personal/characterological attack, addressed to opponent and moderator, no crowd appeal, no frame-breaking |
| C3 Authority-only | low | high | Mild policy content, but plays to the room and names the format as rigged |
| C4 Full package | high | high | Both |
**Held constant across all four:** policy content, candidate name, party cue, apparent age, race, gender, word count (±8%), reading level (±0.5 Flesch-Kincaid grade), and the opponent's preceding turn.
### 2.2 Partisan symmetry — non-negotiable
Every respondent sees the stimulus with a **randomly assigned party cue** (D / R / no-party-stated), crossed with condition. Eight-cell party-by-condition randomization.
This exists because the finding is worthless — and worse, dishonest — if the design can only detect the effect for one side's style. Pre-register the party-symmetry test: if the style effect differs significantly by party cue, that difference is a reported finding, not a footnote. Any writeup that shows the effect for one party must show the other party's estimate at equal prominence.
Two stimulus sets (a Democratic-primary-flavored issue frame and a Republican-primary-flavored one), each rendered in all four conditions, to avoid issue-content confounds with party cue.
### 2.3 Measures
**Dependent variables**
1. Candidate feeling thermometer, 0–100 (primary DV).
2. Vote intention vs. the opponent, forced choice.
3. "Who won this exchange?" — mechanism check.
4. Perceived authenticity, 5-item.
5. Perceived fairness of the format — manipulation check for AUTH.
**Moderators (measured, not assigned)**
- Institutional trust: 4–6 item battery covering media, elections, Congress, courts. Use an existing validated battery rather than writing one.
- Media diet: short-form video hours/week; primary news source; podcast consumption; whether the respondent typically encounters debates as full broadcasts or as clips. That last item is the sleeper variable and should be worded carefully.
- Birth year (continuous — **do not** collect cohort as a category and then treat it as ordinal).
- Political interest, partisanship strength, education, gender, race.
**Manipulation checks:** two items confirming perceived aggression and perceived audience-orientation, used to validate the stimuli in a pilot before full fielding.
### 2.4 Sample and power
Detecting a *difference between two interaction slopes* is expensive. Interaction effects generally require ~4× the N of the corresponding main effect, and H1 is a contrast between two of them.
- **Pilot:** n = 200, purpose is manipulation-check validation only. No hypothesis testing. If the AGGR and AUTH manipulations do not separate cleanly on their checks, revise stimuli and re-pilot.
- **Main field:** n = 2,000, quota-stratified on birth year (four strata, oversampling 1981–1996 and 1997–2005 to ~600 each), party ID, gender, and education.
- Powered for a small-to-moderate interaction contrast (f² ≈ 0.02) at 80%.
- Budget at roughly $1.50–2.00 per respondent for a 6–8 minute instrument: **$3,000–4,000 for the main field**, plus ~$400 for the pilot. If this is out of range, say so now — a 700-person version can test main effects honestly but *cannot* test H1, and running it while claiming otherwise is the exact overclaiming the Parrot audits.
### 2.5 Analysis plan
Primary model, OLS on the thermometer with robust SEs:
```
FT ~ AGGR * TRUST + AUTH * TRUST + FLIP
     + AGGR * AUTH
     + PARTY_CUE * AGGR + PARTY_CUE * AUTH
     + BIRTH_YEAR + covariates
```
**H1 test:** bootstrap the contrast β(AUTH × TRUST) − β(AGGR × TRUST), 10,000 resamples, report the CI on the difference. Pre-register this single contrast as the primary test. Everything else is secondary and labeled as such.
**H3 test:** nested models. Fit cohort × style without trust and media diet, then with. Report the proportion of the cohort interaction coefficient that survives, with a bootstrap CI on the attenuation ratio. Do not use a Sobel test.
**Multiple comparisons:** one primary test. All six per-feature interactions and both party-cue splits are secondary — Benjamini-Hochberg at q = 0.05 across the secondary family, and the family is enumerated in the pre-registration before fielding, not after.
---
## 3. Track A: the observational corpus
Scaffolding and descriptive color. It cannot carry a causal claim and the writeup must not imply that it does.
### 3.1 Corpus
- Primary-debate transcripts, both parties, 2016–2026. Prefer official/network transcripts over auto-captions where available; log the source per document.
- Unit of analysis: the **exchange**, not the article or the full debate.
- Speaker metadata: party, incumbency, polling position at debate date, age, gender, race, insurgent/establishment coding (pre-register the coding rule).
- Debate metadata: network, moderator, stage size, primary vs. general, date.
### 3.2 The normalization trap, again
The 3.4× agentless-passive-by-category finding applies here in a new form: **debate topic drives attack rates.** Immigration and crime segments will run hot on F1/F3 regardless of speaker. Never rank speakers on raw feature rates.
- Topic-code every exchange (fixed taxonomy, ~12 categories).
- Report within-topic comparisons and direct-standardized rates.
- Validate with a mixed-effects model: random intercepts for speaker, debate, and topic.
- **Calibration gate:** before spending LLM tokens on the full corpus, confirm that topic explains a substantial share of raw feature-rate variance on a 300-exchange subsample. If it doesn't, the topic coder is broken — fix it before proceeding.
### 3.3 The clip-selection problem — read this before using engagement as a DV
Engagement-per-view on clipped debate moments is conditioned on the clip having been clipped. Clippers select for exactly the features being measured. This is not a nuisance parameter; it is a mechanism that can generate the entire predicted result out of nothing.
Mitigations, in order of preference:
1. **Preferred:** sample exchanges from the *full transcript*, then check which ones were clipped. Model clipping probability itself as an outcome (`CLIPPED ~ AUTH + AGGR + FLIP + topic + speaker`). This turns the confound into a finding: it measures whether the media ecosystem selects for audience-authority routing, which is a Parrot piece on its own.
2. Restrict engagement analysis to official full-debate uploads where segment-level retention is available.
3. If neither is available, drop the engagement DV entirely. Report feature prevalence descriptively and stop.
Platform age-skew is **not** a cohort measure. If it appears in the writeup at all, it appears with that sentence attached.
### 3.4 Dedup
Same-clip reuploads across accounts are the wire-copy problem in new clothes. Perceptual hash on video, plus transcript n-gram overlap ≥ 0.9 within a 14-day window. One canonical row per underlying exchange, engagement summed with the aggregation rule logged.
---
## 4. Reliability protocol
Matches the existing standard. No rankings, no rates, and no LLM-coded variable enters any model before this gate passes.
- Two human coders independently code a 250-exchange stratified sample.
- **Krippendorff's α ≥ 0.70 per feature**, computed separately for each of F1–F6. A pooled α is not acceptable — F2 and F4 are the hard ones and will hide behind F1's easy agreement.
- Any feature below 0.70: revise the codebook decision rules, re-train, re-code a fresh sample. Two failed rounds on a feature means that feature is dropped from the indices and the drop is disclosed.
- LLM coder is then validated against the human-consensus set. Report LLM-vs-human α per feature alongside human-vs-human. If the LLM underperforms the human floor, it does not get to code the corpus.
- Every LLM code returns a character span. Codes without spans are discarded, not repaired.
---
## 5. Build order
**Phase 0 — Pre-registration.** Hypotheses, the single primary contrast, the enumerated secondary family, kill conditions, stopping rule, exclusion criteria. Timestamped and posted (OSF) before Phase 2 data collection. This is the whole credibility of the project; a study auditing overclaiming that pre-registers after peeking is not recoverable.
**Phase 1 — Codebook + reliability.** §1 and §4. Human coding, α gate, LLM validation. Gate: all retained features ≥ 0.70.
**Phase 2 — Stimulus construction + pilot.** Write eight stimuli (4 conditions × 2 issue frames), verify they differ on the intended dimensions *and only on those*, by human coding with the §1 instrument. Field the n=200 pilot. Gate: manipulation checks separate.
**Phase 3 — Main field.** n = 2,000. Analysis per §2.5. Gate: none — the result is the result.
**Phase 4 — Corpus.** Track A on the Spark. Feature-score the transcript corpus, topic-normalize, run the clipping-probability model. Descriptive only.
**Phase 5 — Writeup.** Both tracks, experiment leading. Publish the pre-registration diff — every deviation from Phase 0, itemized.
---
## 6. Acceptance criteria
1. Pre-registration is public and timestamped before Phase 2 fielding.
2. All retained features clear α ≥ 0.70 for both human-human and LLM-human.
3. Manipulation checks separate AGGR and AUTH in the pilot.
4. Main field hits quota targets within 10% per cohort stratum.
5. The primary contrast is reported with its CI regardless of direction, in the first three paragraphs of any writeup.
6. Party-symmetry estimates are reported at equal prominence.
7. Every observational claim is labeled observational in the sentence that makes it.
8. Deviation log published with results.
---
## 7. What kills this study
- **Stimulus artifact.** If C4 is simply written better or funnier than C1, the study measures writing quality. The pilot's job is to catch this; a humor-rating item in the pilot is cheap insurance.
- **Underpowered contrast.** Running n=700 and reporting H1 anyway.
- **Clip-selection laundering.** Using viral-clip engagement as a proxy for public preference. §3.3 exists to prevent it.
- **Cohort reification.** Reporting "millennials" as an explanation after trust and media diet have eaten the coefficient.
- **Asymmetric framing.** Publishing the party split in only one direction.
- **Post-hoc feature promotion.** Discovering F4 carries everything and rewriting the theory around the flip. If that happens, report it as an exploratory finding requiring independent replication.
---
## 8. Publication paths by outcome
| Outcome | Headline |
|---|---|
| H1 and H2 hold | The debate stage is being scored like a rap battle by low-trust voters — and generation was never the variable |
| H2 fails, aggression carries it | Voters like a fighter. The mechanism is simpler and worse than the elegant version |
| H1 null, both indices positive | Style matters, the components don't separate — a replication call with the instrument attached |
| H3 fails, cohort survives | The generational read holds up under controls that should have killed it |
| Everything null | The instrument, the codebook, and the reliability data published as a tool — plus the clipping-probability model from Phase 4, which is a piece on its own |
Every row is publishable. That is the point of registering it this way.