Palimpsest compares model responses across declared prompts, languages and controls. A screenshot alone cannot establish a repeatable pattern; the evaluation needs retained evidence and explicit limits.
Claim boundaryA controlled comparison can support a scoped claim about named model responses. It cannot establish the behavior of every Chinese model, every political topic, or the motive behind an output.
Each article is built from a closed source record and refuses publication if a cited local artifact is missing. The receipts update when the evidence bytes change.
Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.
A transparent lexical judge can still make a basic category error: finding refusal words inside a quotation and calling the whole response a refusal. Method v4 narrows the decision to the model's own speech act and forces a new baseline.
The first guarded Generative Firewall v2 run published its exact protocol before sampling, retained all 660 sampled responses, and survived a concurrent-main publication race without re-querying.
6 evidence receiptssha256:d291ad260d11
Publication contract
No essay without an exit condition.
Scoped claimWhat this article says—and the broader claim it refuses to make.
Exact evidenceCurrent file path, byte count, and SHA-256 receipt for every local source.
Visible limitsKnown measurement and interpretation boundaries stay next to the conclusion.
FalsifierThe condition that would lower, reverse, or retire the claim.