Live claim ceiling
A green hash is not a validity result
Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.
The most dangerous badge was the easiest one to earn
Hash chains answer a valuable, narrow question: do the served records reproduce their commitments and publication order? They do not answer whether a prompt measures the intended construct, whether a classifier agrees with people, whether a panel represents a larger model population, or whether another team can reproduce the finding.
Collapsing those questions into one ‘verified’ badge rewards the easiest engineering property and launders the hardest scientific gaps. Palimpsest's assurance report refuses that composite. It exposes seven dimensions and assigns each check a status, evidence statement, limitation, affected suite, and—where possible—a local verification command.
Seven questions, no average
Integrity asks whether the registry recomputes and whether every run follows a preregistration. Prompt precommitment asks whether the exact questions were fixed before answers. Response recomputability asks whether full public text can reproduce the current seals. Pipeline reproducibility asks whether labels can be regenerated under an identified method. Statistical design asks for denominators, uncertainty, controls, power, and correction for repeated looks. Construct validation asks whether independent humans support what the labels mean. Replication asks whether an unaffiliated team reproduced the result.
A failure in one dimension is not averaged away by passes elsewhere. The dimension inherits its weakest applicable check, and the public claim ceiling follows the unresolved scientific gates.
- Pass: the served evidence satisfies the declared machine check.
- Partial: a real guarantee exists, but a named part of it is absent or belongs to a legacy method.
- Pending: a preregistered study or evidence-producing step is not complete.
- Open: the invitation exists, but no qualifying independent result is on record.
- Fail: published evidence exists and violates the declared contract or frozen threshold.
The human-validation gate cannot be faked by a filename
The two-coder study is already frozen at 145 rows with a codebook, sample commitment, weighting plan, and three rejection thresholds. Assurance stays pending until one exact result artifact identifies that commitment, accounts for every row from both coders, carries both attestations, and supplies digests that match the released coder sheets, answer key, manifest, and protocol.
Once complete, the check independently recomputes whether Krippendorff's alpha reaches 0.667 and whether weighted precision reaches 0.80 for both refused and party-line labels. A malformed result is fail, not pending. A threshold miss is a published failure, not permission to redraw the sample.
What can be said today
Today Palimpsest can say that its public eval chain is tamper-evident under the declared hash construction; the frontier suite binds exact prompts and exposes full current responses for seal recomputation; and both suites publish explicit denominators, uncertainty, controls, and method versions. The live JSON beside this article reports the exact current count of pass, partial, pending, open, and fail checks.
It cannot yet say that the lexical classifier has completed independent human validation. The legacy China-focused GFI still lacks exact-prompt preregistration and the complete served response matrix that v2 will require. No unaffiliated team has published a preregistered replication against the registry. Those gaps are not footnotes to a strong score. They are the ceiling on the claim.
Why this makes a stronger grant case
A grant should fund a specific reduction in uncertainty, not reward a polished claim that cannot fail. The assurance ladder converts open work into auditable milestones: complete the frozen human study; land the first full-evidence GFI v2 run; recruit and support an unaffiliated replication; extend prompts and coder validation across languages; deposit durable external witnesses; and publish every miss.
That does not guarantee an award. It gives a reviewer something better than ambition: a public baseline, a named gap, a falsifier, a deliverable that changes the claim ceiling only when the evidence earns it, and a verification path that remains after the grant ends.
Limits carried with the claim
What this does not establish
- The assurance report evaluates Palimpsest's declared contracts; it is not an accreditation or an external audit.
- A pass establishes the named machine check only and must be read with that check's limitation.
- Human validation remains pending and independent replication remains open in the current public state.
- The claim ceiling can regress if later evidence fails; promotion is not permanent.
Exit condition
What would change the claim
Any broken registry link, prompt commitment mismatch, unrecomputable response seal, label disagreement under the committed method, missing statistical field, malformed human-study result, frozen-threshold miss, or failed qualifying replication must appear as a non-pass state and lower the affected dimension. If the public page and machine-readable report disagree, the machine-readable check is authoritative and publication should fail.
Related research
Context, not borrowed proof
- NIST: Statistical models expand the AI evaluation toolboxIndependent context on uncertainty in AI evaluation; not an endorsement of Palimpsest
- NIST: Towards best practices for automated benchmark evaluationsIndependent benchmark-governance context; not evidence that Palimpsest satisfies every practice
Reproduce it locally
Verification commands
python -m scripts.verify_eval_registrypython -m scripts.verify_refusal_transcriptspython -m scripts.build_eval_assurance --checkpython -m pytest -q tests/test_eval_assurance.py