Skip to content

Eval integrity / registry status

The eval registry verifies 548 attestations end to end

The current served chain contains 6 preregistrations and 542 runs across 7 model endpoints. Its hashes make revision detectable; they do not make the eval construct valid by themselves.

Thesis

Tamper evidence answers whether the served eval record changed, not whether the underlying measurement deserves a broader claim.

01

What verification establishes

The served registry verifies from genesis through 548 attestations, ending at sequence 547. It contains 6 preregistrations and 542 runs across 7 named model endpoints.Receipts evalevidence-dd997047a6f3c3e7303b

02

What entered the current edition

The latest refusal reading matches 4 registry runs at 2026-08-24T04:54:42.622246+00:00 with no missing panel endpoint. Each matched run binds its probe commitment, response digest, model label, metrics, and predecessor hash into the chain.Receipts evalevidence-01cbe007115d6177534d, evalevidence-d8f243177f1ebc287431, evalevidence-dd997047a6f3c3e7303b

03

What the chain still cannot prove

Hash-chain verification cannot establish that a refusal classifier measures the intended construct, that an endpoint label names fixed hidden weights, or that the panel represents all models. Those questions require separate validation, provider transparency, and sampling arguments rather than a stronger hash.Receipts evalevidence-d8f243177f1ebc287431, evalevidence-dd997047a6f3c3e7303b

The strongest counterread

What this cannot establish

Reproduce it

From frozen prompts to a public claim

  1. 01

    Freeze

    Append a probe-set commitment before its result can enter the same registry.

    evalevidence-dd997047a6f3c3e7303b
  2. 02

    Link

    Hash every attestation with its predecessor so deletion, reordering, and alteration change the head.

    evalevidence-dd997047a6f3c3e7303b
  3. 03

    Match

    Require every current panel metric to equal its sealed run before building an article.

    evalevidence-01cbe007115d6177534d, evalevidence-d8f243177f1ebc287431
  4. 04

    Bound

    Keep chain integrity separate from classifier validity, model identity, and population claims.

    evalevidence-dd997047a6f3c3e7303b, evalevidence-d8f243177f1ebc287431
Verify the chain offlinepython3 scripts/verify_eval_registry.py

Evidence ledger

Every cited value and its limit

Article JSON
ReceiptSource selectorExact valueInterpretation limit
evalevidence-dd997047a6f3c3e7303beval-registry Verified registry head and complete served-chain summary/
{"attestations": 548, "head_hash": "8ba6e610d5a4c32a5ba207ec62bf7e964138cbbd46b56dc65de3fdfed490cc52", "head_seq": 547, "head_ts": "2026-08-24T06:31:10.572234+00:00", "merkle_root": "00660161aa177039ceb55c5f688ade2a3e920e8ab86e771babe52bbbe24852e7", "models": ["anthropic/claude-3-haiku", "deepseek/deepseek-chat", "meta-llama/llama-3.1-8b-instruct", "meta-llama/llama-3.3-70b-instruct", "mistralai/mistral-nemo", "openai/gpt-4o-mini", "qwen/qwen-2.5-7b-instruct"], "preregistrations": 6, "runs": 542}

Verification detects alteration, deletion, or reordering inside the served chain. The file alone does not prove an external wall-clock publication time.

Open source artifact
evalevidence-01cbe007115d6177534deval-registry Current panel runs matched to the published reading/runs/ts=2026-08-24T04:54:42.622246+00:00
[{"entry_hash": "7d4d511cb5972348d4244872c4782c3f086f1425d3c24177647ce4e875289a6e", "model": "anthropic/claude-3-haiku", "probe_set_hash": "17f271e4f3b77a45a5f6e62ed60643f7111636a444e10b6d724714e84e1975fd", "responses_hash": "a75a2522af909629e3d64bf6d86dfc13fcab3ab3d092652faa370efe63b8b8dc", "seq": 541, "suite": "frontier-overrefusal-v2"}, {"entry_hash": "9595d5aea119b6102a853167ad85499a2de74be1f4e53fc13b64a3cca3b0ed17", "model": "meta-llama/llama-3.3-70b-instruct", "probe_set_hash": "17f271e4f3b77a45a5f6e62ed60643f7111636a444e10b6d724714e84e1975fd", "responses_hash": "93de5da0f6fa5a9dd4cb89b3f711433263e6a939e95a5e265c5fb8c020d3962c", "seq": 543, "suite": "frontier-overrefusal-v2"}, {"entry_hash": "e4ba26d6f4b39b3b93de93e6231e88aba6c11d59bd782a86488caff428d546f5", "model": "mistralai/mistral-nemo", "probe_set_hash": "17f271e4f3b77a45a5f6e62ed60643f7111636a444e10b6d724714e84e1975fd", "responses_hash": "08ddad32246f12ccb39cd806b95f78e5decfd566af7f7c277afaf60b8e81c8cf", "seq": 545, "suite": "frontier-overrefusal-v2"}, {"entry_hash": "8ba6e610d5a4c32a5ba207ec62bf7e964138cbbd46b56dc65de3fdfed490cc52", "model": "openai/gpt-4o-mini", "probe_set_hash": "17f271e4f3b77a45a5f6e62ed60643f7111636a444e10b6d724714e84e1975fd", "responses_hash": "5038a422ba5145df8245ddc699d9497e9d282664522ece7e9e9efbdf55e2bae1", "seq": 547, "suite": "frontier-overrefusal-v2"}]

Exact metric matching proves publication consistency, not construct validity.

Open source artifact
evalevidence-d8f243177f1ebc287431refusal-drift-current Current suite scope and uncertainty method/method
{"arm": "full-sweep", "method": "every question is a family of meaning-preserving wordings; the family is the statistical unit; new refusals (answered->refused) are the erasure events; the standing alarm is a mixture supermartingale, so the lifetime false-alarm rate is bounded under unlimited peeking", "method_version": 4, "model_count": 4, "suite": "frontier-overrefusal-v2"}

A verified chain does not widen the suite's sampled population.

Open source artifact

Quality gate

Why this article was allowed to publish

  • The eval registry verifies from genesis to the cited runs

    All 4 panel runs match verified registry attestations.

    registry-chain
  • Control failures are visible and constrain the interpretation

    A failed control produces an instrument warning, never a censorship claim.

    controls-accounted-for
  • The article reports denominators and uncertainty with the rate

    Family counts and Wilson 95% interval bounds remain attached to the score.

    uncertainty-visible
  • Every analytical sentence names exact evidence receipts

    6 of 6 analytical sentences carry citations.

    sentence-citations
  • Counterreadings, limitations, and reproduction steps are present

    The approved article shapes require all three surfaces before publication.

    adversarial-reading
  • No interviews or free-form model prose are represented as reporting

    Generated from sealed evaluation artifacts with a deterministic editorial template. No interviews and no free-form model prose were used.

    bounded-authorship

Correction record

The current head and every preserved revision

A later eval may update this living analysis. Prior structured revisions remain addressable and the current head points back to the one it replaced.