PALIMPSEST / LIVE EVAL FINDINGSEdition 11 Aug 2026, 02:28 UTC
Interpretation with receipts attached
The Eval Journal
AI evaluation scores are easy to publish and hard to read honestly. This desk starts with the controls, carries uncertainty into the headline, and binds every analytical sentence to a sealed run.
one score, one ranking, one easy answerdated finding, visible limits, reproducible record
Meta Llama 3.3 70B Instruct passed all 144 prompt arms in the latest full-sweep run. The prior failed controls remain part of the record; the change is descriptive rather than causal.
4 control-clean model runs observed 0 refused families out of 34, yet each zero result still carries a Wilson 95% upper bound of 10.2%. That interval is part of the finding, not fine print.
Across 4 sealed endpoint runs, 3 of 163 paired family comparisons changed state: 0 toward refusal and 3 toward answer. This is a dated transition, not a trend claim.
The current served chain contains 7 preregistrations and 620 runs across 7 model endpoints. Its hashes make revision detectable; they do not make the eval construct valid by themselves.