Skip to content
The Jev trial: offline test results and trial terms
offline eval and shadow-trial terms, run on the DGX by the desk's Claude Code session (claude-opus-5-5) · 1 turns · 2026-09-29 06:00–08:00 UTC · body sha256 fa9999e382ac · results file ~/jobs/jev-span-eval/results/SUMMARY.md, reproduced verbatim
The Jev trial: offline test results and trial terms
Offline test run on the desk DGX, finished 2026-09-28 23:36 PDT. Script: ~/jobs/jev-span-eval/jev_eval.py. Reads archived runs only. Model pinned: jev-1.13.0.
# Jev offline eval (jev-1.13.0)
Spend: $0.0209 over 554 live calls (496,700 input tokens); cached answers cost nothing.
## A. Rescuing damaged spans
- Jev picked the right sentence: 148/150 (99%)
- Fuzzy-match baseline: 132/150 (88%)
- Jev said none on a real (damaged) span: 2/150 (1%)
- At confidence >= 0.8: 124/124 (100%) correct, covering 124/150 (83%) of cases
| damage | Jev | baseline |
| --- | --- | --- |
| paraphrase | 40/40 (100%) | 34/40 (85%) |
| truncate | 39/39 (100%) | 33/39 (85%) |
| glyph | 30/32 (94%) | 28/32 (88%) |
| merge | 39/39 (100%) | 37/39 (95%) |
**Negative controls** (quote from a different article; right answer is none): Jev said none on 50/50 (100%); false rescues at confidence >= 0.8: 0.
Baseline best-match ratio on controls ranges up to 0.456; on real cases the median is 0.658. A threshold between them decides how many false rescues the baseline would make.
## B. Agreement with N7's discrepancy verdicts
Overall agreement: 128/359 (36%) (N7 is the current pipeline, not ground truth)
| N7 said \ Jev said | hard_contradiction | naming_inconsistency | framing_divergence | compatible |
| --- | --- | --- | --- | --- |
| hard_contradiction | 106 | 3 | 20 | 107 |
| naming_inconsistency | 2 | 7 | 11 | 40 |
| framing_divergence | 5 | 0 | 12 | 43 |
| compatible | 0 | 0 | 0 | 3 |
- Of N7's hard contradictions, Jev also called hard: 106/236 (45%)
- Of Jev's hard calls, N7 agreed: 106/113 (94%)
Classification sample drawn from the archive:
[B] sample: Counter({'hard_contradiction': 236, 'naming_inconsistency': 60, 'framing_divergence': 60, 'compatible': 3})
Trial terms (parrot/jev_shadow.py, commit 6de8ab7b4, switched on 2026-09-28 23:43 PDT):
- Runs after N14 on every shift when JEV_SHADOW=1; appends to data/operator/jev_shadow.jsonl; nothing in the pipeline reads it back.
- Model pinned: jev-1.13.0.
- Daily spend cap: JEV_SHADOW_DAILY_USD=0.25.
- Per-shift caps: 60 calls, 90 seconds.
- Live span rescue only after at least 30 real rescue cases are logged and checked; a rescued span is accepted only at confidence >= 0.8, otherwise dropped as today.
- Contradiction labels stay shadow-only until disagreements are read against the record.
- A search of data/logs and data/operator on 2026-09-29 for historical per-shift unlocatable-span counts found none.