Skip to content

Eval method / uncertainty

Zero observed refusals is not zero uncertainty

4 control-clean model runs observed 0 refused families out of 34, yet each zero result still carries a Wilson 95% upper bound of 10.2%. That interval is part of the finding, not fine print.

Thesis

A finite eval can observe no refusals and still leave a meaningful range of plausible rates.

01

What zero contains

In the 11 Aug 2026 panel, 4 control-clean runs observed zero refused families among 34 monitored non-control families. For a zero-of-34 result, the published Wilson 95% interval still reaches 10.2%. Zero observed events and zero plausible event rate are different statements.Receipts evalevidence-e00da8ff59ae037434fc, evalevidence-e4d8af3496e0ac6acdd0, evalevidence-e9d19112ab67a6bbd3c4

02

The unit is a question family, not a prompt

Each monitored question can appear in several meaning-preserving wordings, but the family is the statistical unit. That prevents a model from looking artificially precise merely because the same idea was phrased many times.Receipts evalevidence-e00da8ff59ae037434fc

03

Cross-lab agreement is useful and bounded

The latest panel contains 4 named endpoints, of which 4 passed every ordinary control. Agreement across those endpoints is stronger than a one-model anecdote, but it is not a probability sample of all models, deployments, languages, or future releases.Receipts evalevidence-e00da8ff59ae037434fc, evalevidence-e4d8af3496e0ac6acdd0

04

Read it as a dated panel, not a leaderboard

The registry binds every score to a model label, prompt commitment, response hash, and timestamp. Those receipts make change auditable, but they do not justify a permanent ranking from one sweep.Receipts evalevidence-e4d8af3496e0ac6acdd0, evalevidence-e9d19112ab67a6bbd3c4

The strongest counterread

What this cannot establish

Reproduce it

From frozen prompts to a public claim

  1. 01

    Pre-register

    Commit the exact probe bank before any endpoint is queried.

    evalevidence-e9d19112ab67a6bbd3c4
  2. 02

    Group

    Treat a meaning-preserving prompt family as the statistical unit.

    evalevidence-e00da8ff59ae037434fc
  3. 03

    Bound

    Report the denominator and Wilson 95% interval with every family rate.

    evalevidence-e4d8af3496e0ac6acdd0, evalevidence-e00da8ff59ae037434fc
  4. 04

    Re-run

    Compare later sealed sweeps with the same method instead of turning this edition into a standing ranking.

    evalevidence-e9d19112ab67a6bbd3c4
Verify the chain offlinepython3 scripts/verify_eval_registry.py

Evidence ledger

Every cited value and its limit

Article JSON
ReceiptSource selectorExact valueInterpretation limit
evalevidence-e4d8af3496e0ac6acdd0refusal-drift-current Latest cross-lab family-level panel/models/*/{family_refusal_rate_pct,family_refusal_ci95_pct,n_families,controls_clean}
{"anthropic/claude-3-haiku": {"controls_clean": true, "family_refusal_ci95_pct": [0.0, 10.2], "family_refusal_rate_pct": 0.0, "n_families": 34, "n_refused_families": 0}, "meta-llama/llama-3.3-70b-instruct": {"controls_clean": true, "family_refusal_ci95_pct": [0.0, 10.2], "family_refusal_rate_pct": 0.0, "n_families": 34, "n_refused_families": 0}, "mistralai/mistral-nemo": {"controls_clean": true, "family_refusal_ci95_pct": [0.0, 10.2], "family_refusal_rate_pct": 0.0, "n_families": 34, "n_refused_families": 0}, "openai/gpt-4o-mini": {"controls_clean": true, "family_refusal_ci95_pct": [0.0, 10.2], "family_refusal_rate_pct": 0.0, "n_families": 34, "n_refused_families": 0}}

The panel is a dated, non-representative set of named model endpoints.

Open source artifact
evalevidence-e00da8ff59ae037434fcrefusal-drift-current Family-level statistical method/{method,method_note,n_families,control_families}
{"control_families": ["civics/how-a-bill", "control/bread-baking", "control/kettle-limescale", "control/offside-rule", "control/photosynthesis", "control/ser-estar", "control/thunder-delay"], "method": "every question is a family of meaning-preserving wordings; the family is the statistical unit; new refusals (answered->refused) are the erasure events; the standing alarm is a mixture supermartingale, so the lifetime false-alarm rate is bounded under unlimited peeking", "method_note": "Rates are family-level with Wilson 95% intervals. The paired test is an exact mid-p McNemar on a single transition and is not valid for the rolling series; the churn monitor is, by Ville's inequality. Control families are unremarkable questions: if they are refused, the run is an instrument fault and carries no censorship claim. The refusal classifier is lexical and is itself watched by a frozen anchor set.", "n_families": 41}

Wilson intervals quantify sampling uncertainty inside the suite, not deployment-wide uncertainty.

Open source artifact
evalevidence-e9d19112ab67a6bbd3c4eval-registry Sealed attestations for the latest panel/runs/@ts=2026-08-11T02:28:16.121324+00:00/@suite=frontier-overrefusal-v2
{"anthropic/claude-3-haiku": {"entry_hash": "62a21b26a246a63eb78ad1c01b43fa94d0f98961c26086ef0bdc77b62903ffd8", "responses_hash": "e0e792c2253bb3bf533a7bbade0856219c21da9acc4aa0dc19a40b5dcb1df0c8", "seq": 304}, "meta-llama/llama-3.3-70b-instruct": {"entry_hash": "0320438ee17a3f10c1cf0de23faf50756352c90456bca1a308a5172aaaabb9bd", "responses_hash": "c3608f9ef26c459ae843bbc6fb9ed4ba9e5be8f2d14e2f9cf898a5098fd4be75", "seq": 306}, "mistralai/mistral-nemo": {"entry_hash": "fedd2360e36b61f3fc13abd4ee4b7a13442a355476e16a7f46838aa4d741b7ea", "responses_hash": "ce0ba5f725ece9103ef2ba9ac215fc6232dfa3d379c398cd5cb190df6adb566a", "seq": 308}, "openai/gpt-4o-mini": {"entry_hash": "a1197f007cbf1fdee8b5ae9b17b79aa0c843f6e7193aa55c28647d1ddad4ede4", "responses_hash": "9c0c29fd1f190a61cc9d1d4a54b2c9894958960ed652e4be0aa5534cb67e2827", "seq": 302}}

The chain proves these attestations persisted unchanged. It does not widen the sampled population.

Open source artifact

Quality gate

Why this article was allowed to publish

  • The eval registry verifies from genesis to the cited runs

    All 4 panel runs match verified registry attestations.

    registry-chain
  • Control failures are visible and constrain the interpretation

    A failed control produces an instrument warning, never a censorship claim.

    controls-accounted-for
  • The article reports denominators and uncertainty with the rate

    Family counts and Wilson 95% interval bounds remain attached to the score.

    uncertainty-visible
  • Every analytical sentence names exact evidence receipts

    9 of 9 analytical sentences carry citations.

    sentence-citations
  • Counterreadings, limitations, and reproduction steps are present

    The approved article shapes require all three surfaces before publication.

    adversarial-reading
  • No interviews or free-form model prose are represented as reporting

    Generated from sealed evaluation artifacts with a deterministic editorial template. No interviews and no free-form model prose were used.

    bounded-authorship

Correction record

The current head and every preserved revision

A later eval may update this living analysis. Prior structured revisions remain addressable and the current head points back to the one it replaced.