PALIMPSEST / LIVE EVAL FINDINGSEdition 24 Aug 2026, 04:54 UTC
Interpretation with receipts attached
The Eval Journal
AI evaluation scores are easy to publish and hard to read honestly. This desk starts with the controls, carries uncertainty into the headline, and binds every analytical sentence to a sealed run.
one score, one ranking, one easy answerdated finding, visible limits, reproducible record
Meta Llama 3.3 70B Instruct refused 0 of 34 monitored question families, but 1 ordinary control prompt arms also refused. That makes the result an instrument warning, not a censorship finding.
3 control-clean model runs observed 0 refused families out of 34, yet each zero result still carries a Wilson 95% upper bound of 10.2%. That interval is part of the finding, not fine print.
Across 4 sealed endpoint runs, 17 of 162 paired family comparisons changed state: 0 toward refusal and 17 toward answer. This is a dated transition, not a trend claim. 1 run also failed ordinary controls.
The current served chain contains 6 preregistrations and 542 runs across 7 model endpoints. Its hashes make revision detectable; they do not make the eval construct valid by themselves.