{
  "author": "Palimpsest Eval Lab",
  "claim": "The current eval record supports a provisional measurement claim; it does not yet support calling the lexical construct human-validated or the findings independently replicated.",
  "content_sha256": "481deb73f5ae136a08961ce138dd544cfdde949f83a0e79943db0c80b202c8af",
  "dek": "Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.",
  "evidence": [
    {
      "bytes": 9261,
      "label": "AI Eval Assurance",
      "path": "readings/eval-assurance-latest.json",
      "role": "Live dimensions, checks, claim ceiling, limitations, and verification commands",
      "sha256": "9b02ed1d778a38e41b7a717671272ed54f057bce6c1d2abd0f7815dbab664f38",
      "url": "/readings/eval-assurance-latest.json"
    },
    {
      "bytes": 29615,
      "label": "Assurance builder",
      "path": "core/eval_assurance.py",
      "role": "Deterministic checks and human-validation promotion gate",
      "sha256": "a3bad38968408bae2028d921210fe1397c09b2c5b00b6c2e309aac0828f26b1d",
      "url": "/core/eval_assurance.py"
    },
    {
      "bytes": 2868,
      "label": "Closed assurance schema",
      "path": "protocol/eval-assurance-v1.schema.json",
      "role": "Machine contract for the served report",
      "sha256": "aeb8f49e3e8aba2c570085224c84b50e8af72c39c4845139f3fa303443d61b25",
      "url": "/protocol/eval-assurance-v1.schema.json"
    },
    {
      "bytes": 4768,
      "label": "Frozen validation study",
      "path": "validation/studies/2026-08-01-gfi-classifier-v1/PROTOCOL.json",
      "role": "Sample, analysis, weighting, thresholds, and one-look commitment",
      "sha256": "1cd6b00594a8a737c0604ec4e86f7a25d3c5ddeb72a909c2c0b7a74576db8cbd",
      "url": "/validation/studies/2026-08-01-gfi-classifier-v1/PROTOCOL.json"
    },
    {
      "bytes": 11042,
      "label": "Grant evidence case",
      "path": "docs/GRANT-CASE.md",
      "role": "Fundable work packages, milestones, risks, and non-negotiable claim boundaries",
      "sha256": "e2de095d71d225e469ec25910671032d9e2fb7a8d13bb945fe202a2f3b294134",
      "url": "/docs/GRANT-CASE.md"
    },
    {
      "bytes": 3202,
      "label": "Assurance regression tests",
      "path": "tests/test_eval_assurance.py",
      "role": "Pins chain failure, filename spoofing, malformed-result failure, and current state",
      "sha256": "6de4672516b59178dc86a84106774c55c9b48dbe2364ae7fe8e4cdf8e03200a3",
      "url": "/tests/test_eval_assurance.py"
    }
  ],
  "external_sources": [
    {
      "relationship": "Independent context on uncertainty in AI evaluation; not an endorsement of Palimpsest",
      "title": "NIST: Statistical models expand the AI evaluation toolbox",
      "url": "https://www.nist.gov/news-events/news/2026/02/new-report-expanding-ai-evaluation-toolbox-statistical-models"
    },
    {
      "relationship": "Independent benchmark-governance context; not evidence that Palimpsest satisfies every practice",
      "title": "NIST: Towards best practices for automated benchmark evaluations",
      "url": "https://www.nist.gov/news-events/news/2026/01/towards-best-practices-automated-benchmark-evaluations"
    }
  ],
  "falsifier": "Any broken registry link, prompt commitment mismatch, unrecomputable response seal, label disagreement under the committed method, missing statistical field, malformed human-study result, frozen-threshold miss, or failed qualifying replication must appear as a non-pass state and lower the affected dimension. If the public page and machine-readable report disagree, the machine-readable check is authoritative and publication should fail.",
  "json_url": "https://palimpsest.info/evals/what-the-evidence-can-claim/article.json",
  "kind": "Assurance note",
  "limitations": [
    "The assurance report evaluates Palimpsest's declared contracts; it is not an accreditation or an external audit.",
    "A pass establishes the named machine check only and must be read with that check's limitation.",
    "Human validation remains pending and independent replication remains open in the current public state.",
    "The claim ceiling can regress if later evidence fails; promotion is not permanent."
  ],
  "live_context": {
    "detail": "9 pass · 0 partial · 1 pending · 1 open · 0 fail",
    "label": "Live assurance ceiling",
    "url": "/readings/eval-assurance-latest.json",
    "value": "provisional measurement"
  },
  "modified_at": "2026-08-24T06:31:10.572234Z",
  "published_at": "2026-08-14T17:54:00Z",
  "schema": "palimpsest.eval-journal-article.v1",
  "sections": [
    {
      "heading": "The most dangerous badge was the easiest one to earn",
      "paragraphs": [
        "Hash chains answer a valuable, narrow question: do the served records reproduce their commitments and publication order? They do not answer whether a prompt measures the intended construct, whether a classifier agrees with people, whether a panel represents a larger model population, or whether another team can reproduce the finding.",
        "Collapsing those questions into one ‘verified’ badge rewards the easiest engineering property and launders the hardest scientific gaps. Palimpsest's assurance report refuses that composite. It exposes seven dimensions and assigns each check a status, evidence statement, limitation, affected suite, and—where possible—a local verification command."
      ],
      "points": []
    },
    {
      "heading": "Seven questions, no average",
      "paragraphs": [
        "Integrity asks whether the registry recomputes and whether every run follows a preregistration. Prompt precommitment asks whether the exact questions were fixed before answers. Response recomputability asks whether full public text can reproduce the current seals. Pipeline reproducibility asks whether labels can be regenerated under an identified method. Statistical design asks for denominators, uncertainty, controls, power, and correction for repeated looks. Construct validation asks whether independent humans support what the labels mean. Replication asks whether an unaffiliated team reproduced the result.",
        "A failure in one dimension is not averaged away by passes elsewhere. The dimension inherits its weakest applicable check, and the public claim ceiling follows the unresolved scientific gates."
      ],
      "points": [
        "Pass: the served evidence satisfies the declared machine check.",
        "Partial: a real guarantee exists, but a named part of it is absent or belongs to a legacy method.",
        "Pending: a preregistered study or evidence-producing step is not complete.",
        "Open: the invitation exists, but no qualifying independent result is on record.",
        "Fail: published evidence exists and violates the declared contract or frozen threshold."
      ]
    },
    {
      "heading": "The human-validation gate cannot be faked by a filename",
      "paragraphs": [
        "The two-coder study is already frozen at 145 rows with a codebook, sample commitment, weighting plan, and three rejection thresholds. Assurance stays pending until one exact result artifact identifies that commitment, accounts for every row from both coders, carries both attestations, and supplies digests that match the released coder sheets, answer key, manifest, and protocol.",
        "Once complete, the check independently recomputes whether Krippendorff's alpha reaches 0.667 and whether weighted precision reaches 0.80 for both refused and party-line labels. A malformed result is fail, not pending. A threshold miss is a published failure, not permission to redraw the sample."
      ],
      "points": []
    },
    {
      "heading": "What can be said today",
      "paragraphs": [
        "Today Palimpsest can say that its public eval chain is tamper-evident under the declared hash construction; the frontier suite binds exact prompts and exposes full current responses for seal recomputation; and both suites publish explicit denominators, uncertainty, controls, and method versions. The live JSON beside this article reports the exact current count of pass, partial, pending, open, and fail checks.",
        "It cannot yet say that the lexical classifier has completed independent human validation. The legacy China-focused GFI still lacks exact-prompt preregistration and the complete served response matrix that v2 will require. No unaffiliated team has published a preregistered replication against the registry. Those gaps are not footnotes to a strong score. They are the ceiling on the claim."
      ],
      "points": []
    },
    {
      "heading": "Why this makes a stronger grant case",
      "paragraphs": [
        "A grant should fund a specific reduction in uncertainty, not reward a polished claim that cannot fail. The assurance ladder converts open work into auditable milestones: complete the frozen human study; land the first full-evidence GFI v2 run; recruit and support an unaffiliated replication; extend prompts and coder validation across languages; deposit durable external witnesses; and publish every miss.",
        "That does not guarantee an award. It gives a reviewer something better than ambition: a public baseline, a named gap, a falsifier, a deliverable that changes the claim ceiling only when the evidence earns it, and a verification path that remains after the grant ends."
      ],
      "points": []
    }
  ],
  "slug": "what-the-evidence-can-claim",
  "status": "Live claim ceiling",
  "title": "A green hash is not a validity result",
  "updated_at": "2026-08-14T19:15:00Z",
  "url": "https://palimpsest.info/evals/what-the-evidence-can-claim/",
  "verification": [
    "python -m scripts.verify_eval_registry",
    "python -m scripts.verify_refusal_transcripts",
    "python -m scripts.build_eval_assurance --check",
    "python -m pytest -q tests/test_eval_assurance.py"
  ]
}
