Skip to content

Palimpsest / AI Eval Journal

Edition 001
4 evidence-bound essays
Updated 08 October 2026 · 08:23 UTC

Methods, failures, and findings from the lab measuring how language models refuse, erase, or reframe contested information.

01Prompt

Freeze the exact question and comparison before the answer exists.

02Discrepancy

Measure refusal, substitution, asymmetry, controls, and uncertainty.

03Proof

Publish the response bytes, seals, limits, and the test that could fail.

Method note · Methods and claim limits

A censored answer is not yet evidence

Palimpsest compares model responses across declared prompts, languages and controls. A screenshot alone cannot establish a repeatable pattern; the evaluation needs retained evidence and explicit limits.

Claim boundaryA controlled comparison can support a scoped claim about named model responses. It cannot establish the behavior of every Chinese model, every political topic, or the motive behind an output.

Read the evaluation method →

Launch file

What changed in the eval engine

Each article is built from a closed source record and refuses publication if a cited local artifact is missing. The receipts update when the evidence bytes change.

Assurance note · Live claim ceiling

A green hash is not a validity result

Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.

6 evidence receipts sha256:e8f4ca395119

Method autopsy · Judge v4 shipped; rebaseline pending

When ‘I cannot help’ is evidence of an answer

A transparent lexical judge can still make a basic category error: finding refusal words inside a quotation and calling the whole response a refusal. Method v4 narrows the decision to the model's own speech act and forces a new baseline.

6 evidence receipts sha256:5b5b704b05ba

Protocol note · First sealed v2 run live

GFI v2: the answer comes after the protocol

The first guarded Generative Firewall v2 run published its exact protocol before sampling, retained all 660 sampled responses, and survived a concurrent-main publication race without re-querying.

6 evidence receipts sha256:d291ad260d11

Publication contract

No essay without an exit condition.

  1. Scoped claimWhat this article says—and the broader claim it refuses to make.
  2. Exact evidenceCurrent file path, byte count, and SHA-256 receipt for every local source.
  3. Visible limitsKnown measurement and interpretation boundaries stay next to the conclusion.
  4. FalsifierThe condition that would lower, reverse, or retire the claim.
Follow the methods deskRSSJSON FeedStructured edition