Methods and claim limits
A censored answer is not yet evidence
Palimpsest compares model responses across declared prompts, languages and controls. A screenshot alone cannot establish a repeatable pattern; the evaluation needs retained evidence and explicit limits.
What the evaluation measures
Palimpsest tests Chinese and state-aligned language models with declared prompts about documented political events and criticism of the Chinese Communist Party. The comparison distinguishes refusal, omission and narrative substitution across the named models and languages.
A screenshot preserves a single response. It cannot establish whether the prompt was selected after the answer, whether sampling changed the result, whether a safety refusal was misclassified, or whether the endpoint changed. Those questions require a declared comparison and retained response evidence.
What the record contains
Palimpsest compares declared models, languages, prompt families and neutral controls. Unreachable calls remain abstentions. Each result carries denominators and uncertainty; protocol commitments and retained response evidence support checks on collection order and later revision.
These records let a reviewer inspect selection, omission and method changes. Their integrity does not establish that the classifier is valid or that a result generalizes beyond the tested panel.
- Prompt: freeze the question, language, model panel, sampling plan, and judge before collection.
- Discrepancy: compare refusal, narrative substitution, language asymmetry, prompt sensitivity, and controls without treating one model as ground truth for another.
- Proof: publish the response bytes, hashes, denominators, method version, verification commands, and the limits on the claim.
How model evaluations relate to the observatory
Model responses are a separate measurement layer alongside DNS interference, content deletion, search suppression and storefront removal. A model answer is not averaged into a national censorship score.
The surrounding observatory supplies context and possible comparisons. It does not establish causation, and a refusal does not establish a government instruction. Palimpsest measures named endpoints on named dates under a declared prompt bank.
When the claim must narrow
Failed controls require abstention. A classifier change requires a new baseline. Human coding and independent replication must be evaluated separately from a valid run seal, and negative findings remain part of the record.
The published claim must narrow when repeated samples, matched language pairs or independent checks fail to support it. The question is whether another reviewer can check the result against the declared test and retained evidence.
Limits carried with the claim
What this does not establish
- This method note describes the evaluation design; it is not itself an experimental result.
- Palimpsest evaluates named model endpoints, prompts, and dates. It does not estimate all Chinese models or all political knowledge.
- Observed refusal or narrative substitution does not identify who caused the behavior or prove a government instruction.
- The current lexical labels remain provisional until the preregistered two-human study completes.
Exit condition
What would change the claim
If preregistered comparisons with clean controls, repeated samples, prompt families and language pairs do not show a stable discrepancy on the named model endpoints, the claim does not generalize. If human coders reject the labels or an unaffiliated replication fails, the corresponding claim must narrow; the test must not be redrawn after the result.
Related research
Context, not borrowed proof
- Refusal and reframing in China-origin vision-language modelsAdjacent independent research; not evidence for any Palimpsest measurement
- Bilingual political-bias evaluation around TaiwanRelated evidence that language and political topic can interact; not a replication of Palimpsest
Reproduce it locally
Verification commands
python -m scripts.verify_eval_registrypython -m scripts.verify_refusal_transcriptspython -m scripts.build_eval_assurance --check