Skip to content
PALIMPSEST / LIVE EVAL FINDINGS Edition 24 Aug 2026, 04:54 UTC

Interpretation with receipts attached

The Eval
Journal

AI evaluation scores are easy to publish and hard to read honestly. This desk starts with the controls, carries uncertainty into the headline, and binds every analytical sentence to a sealed run.

04current analyses

Latest from the desk

Read the finding. Then inspect the instrument.

Follow every verified revision

Eval integrity / control failure

The headline was 0%. The controls still failed.

Meta Llama 3.3 70B Instruct refused 0 of 34 monitored question families, but 1 ordinary control prompt arms also refused. That makes the result an instrument warning, not a censorship finding.

  • 0%family refusal rate
  • 1control prompt refusals
  • 3/4panel runs with clean controls

Palimpsest Eval Desk

Read the analysis

Eval method / uncertainty

Zero observed refusals is not zero uncertainty

3 control-clean model runs observed 0 refused families out of 34, yet each zero result still carries a Wilson 95% upper bound of 10.2%. That interval is part of the finding, not fine print.

  • 0refused families
  • 34monitored non-control families
  • 10.2%95% upper interval bound

Palimpsest Eval Desk

Read the analysis

Model behavior / adjacent-run drift

17 previously refused answers returned in the latest panel

Across 4 sealed endpoint runs, 17 of 162 paired family comparisons changed state: 0 toward refusal and 17 toward answer. This is a dated transition, not a trend claim. 1 run also failed ordinary controls.

  • 17answer-state transitions
  • 162paired family comparisons
  • 1/4endpoints with a change

Palimpsest Eval Desk

Read the analysis

Eval integrity / registry status

The eval registry verifies 548 attestations end to end

The current served chain contains 6 preregistrations and 542 runs across 7 model endpoints. Its hashes make revision detectable; they do not make the eval construct valid by themselves.

  • 548verified attestations
  • 6preregistrations
  • 542sealed runs

Palimpsest Eval Desk

Read the analysis

Publication standard

The eval does not get the last word.

  1. 01

    Seal first

    The prompt commitment and result are verified against the append-only registry before a sentence is written.

  2. 02

    Controls before score

    A failed ordinary control turns the article into an instrument warning. It cannot become a suppression claim.

  3. 03

    Uncertainty in the headline

    Rates keep their denominator, interval, model label, and timestamp. Nothing becomes a standing leaderboard.

  4. 04

    Show the counterread

    Every piece carries the strongest alternative reading, limitations, and the command needed to verify the chain.

Follow the record The prose is a route into the evidence, never a replacement for it.