{
  "authors": [
    {
      "name": "Palimpsest Eval Lab",
      "url": "https://palimpsest.info/evals/"
    }
  ],
  "description": "Evidence-bound essays about censorship evaluations, method changes, failures, and what the current record can actually support.",
  "feed_url": "https://palimpsest.info/evals/feed.json",
  "home_page_url": "https://palimpsest.info/evals/",
  "items": [
    {
      "_palimpsest": {
        "article_json": "https://palimpsest.info/evals/what-the-evidence-can-claim/article.json",
        "claim": "The current eval record supports a provisional measurement claim; it does not yet support calling the lexical construct human-validated or the findings independently replicated.",
        "content_sha256": "481deb73f5ae136a08961ce138dd544cfdde949f83a0e79943db0c80b202c8af",
        "falsifier": "Any broken registry link, prompt commitment mismatch, unrecomputable response seal, label disagreement under the committed method, missing statistical field, malformed human-study result, frozen-threshold miss, or failed qualifying replication must appear as a non-pass state and lower the affected dimension. If the public page and machine-readable report disagree, the machine-readable check is authoritative and publication should fail.",
        "kind": "eval_method_article",
        "schema": "palimpsest.eval-journal-article.v1"
      },
      "attachments": [
        {
          "mime_type": "application/json",
          "size_in_bytes": 9261,
          "title": "AI Eval Assurance",
          "url": "https://palimpsest.info/readings/eval-assurance-latest.json"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 29615,
          "title": "Assurance builder",
          "url": "https://palimpsest.info/core/eval_assurance.py"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 2868,
          "title": "Closed assurance schema",
          "url": "https://palimpsest.info/protocol/eval-assurance-v1.schema.json"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 4768,
          "title": "Frozen validation study",
          "url": "https://palimpsest.info/validation/studies/2026-08-01-gfi-classifier-v1/PROTOCOL.json"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 11042,
          "title": "Grant evidence case",
          "url": "https://palimpsest.info/docs/GRANT-CASE.md"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 3202,
          "title": "Assurance regression tests",
          "url": "https://palimpsest.info/tests/test_eval_assurance.py"
        }
      ],
      "authors": [
        {
          "name": "Palimpsest Eval Lab"
        }
      ],
      "content_text": "A green hash is not a validity result\n\nPalimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.\n\nClaim boundary: The current eval record supports a provisional measurement claim; it does not yet support calling the lexical construct human-validated or the findings independently replicated.\n\nThe most dangerous badge was the easiest one to earn\n\nHash chains answer a valuable, narrow question: do the served records reproduce their commitments and publication order? They do not answer whether a prompt measures the intended construct, whether a classifier agrees with people, whether a panel represents a larger model population, or whether another team can reproduce the finding.\n\nCollapsing those questions into one ‘verified’ badge rewards the easiest engineering property and launders the hardest scientific gaps. Palimpsest's assurance report refuses that composite. It exposes seven dimensions and assigns each check a status, evidence statement, limitation, affected suite, and—where possible—a local verification command.\n\nSeven questions, no average\n\nIntegrity asks whether the registry recomputes and whether every run follows a preregistration. Prompt precommitment asks whether the exact questions were fixed before answers. Response recomputability asks whether full public text can reproduce the current seals. Pipeline reproducibility asks whether labels can be regenerated under an identified method. Statistical design asks for denominators, uncertainty, controls, power, and correction for repeated looks. Construct validation asks whether independent humans support what the labels mean. Replication asks whether an unaffiliated team reproduced the result.\n\nA failure in one dimension is not averaged away by passes elsewhere. The dimension inherits its weakest applicable check, and the public claim ceiling follows the unresolved scientific gates.\n\n- Pass: the served evidence satisfies the declared machine check.\n\n- Partial: a real guarantee exists, but a named part of it is absent or belongs to a legacy method.\n\n- Pending: a preregistered study or evidence-producing step is not complete.\n\n- Open: the invitation exists, but no qualifying independent result is on record.\n\n- Fail: published evidence exists and violates the declared contract or frozen threshold.\n\nThe human-validation gate cannot be faked by a filename\n\nThe two-coder study is already frozen at 145 rows with a codebook, sample commitment, weighting plan, and three rejection thresholds. Assurance stays pending until one exact result artifact identifies that commitment, accounts for every row from both coders, carries both attestations, and supplies digests that match the released coder sheets, answer key, manifest, and protocol.\n\nOnce complete, the check independently recomputes whether Krippendorff's alpha reaches 0.667 and whether weighted precision reaches 0.80 for both refused and party-line labels. A malformed result is fail, not pending. A threshold miss is a published failure, not permission to redraw the sample.\n\nWhat can be said today\n\nToday Palimpsest can say that its public eval chain is tamper-evident under the declared hash construction; the frontier suite binds exact prompts and exposes full current responses for seal recomputation; and both suites publish explicit denominators, uncertainty, controls, and method versions. The live JSON beside this article reports the exact current count of pass, partial, pending, open, and fail checks.\n\nIt cannot yet say that the lexical classifier has completed independent human validation. The legacy China-focused GFI still lacks exact-prompt preregistration and the complete served response matrix that v2 will require. No unaffiliated team has published a preregistered replication against the registry. Those gaps are not footnotes to a strong score. They are the ceiling on the claim.\n\nWhy this makes a stronger grant case\n\nA grant should fund a specific reduction in uncertainty, not reward a polished claim that cannot fail. The assurance ladder converts open work into auditable milestones: complete the frozen human study; land the first full-evidence GFI v2 run; recruit and support an unaffiliated replication; extend prompts and coder validation across languages; deposit durable external witnesses; and publish every miss.\n\nThat does not guarantee an award. It gives a reviewer something better than ambition: a public baseline, a named gap, a falsifier, a deliverable that changes the claim ceiling only when the evidence earns it, and a verification path that remains after the grant ends.\n\nLimitations\n\n- The assurance report evaluates Palimpsest's declared contracts; it is not an accreditation or an external audit.\n\n- A pass establishes the named machine check only and must be read with that check's limitation.\n\n- Human validation remains pending and independent replication remains open in the current public state.\n\n- The claim ceiling can regress if later evidence fails; promotion is not permanent.\n\nWhat would change the claim\n\nAny broken registry link, prompt commitment mismatch, unrecomputable response seal, label disagreement under the committed method, missing statistical field, malformed human-study result, frozen-threshold miss, or failed qualifying replication must appear as a non-pass state and lower the affected dimension. If the public page and machine-readable report disagree, the machine-readable check is authoritative and publication should fail.",
      "date_modified": "2026-08-24T06:31:10.572234Z",
      "date_published": "2026-08-14T17:54:00Z",
      "id": "481deb73f5ae136a08961ce138dd544cfdde949f83a0e79943db0c80b202c8af",
      "summary": "Palimpsest evaluation method article. Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.",
      "tags": [
        "AI evaluations",
        "Assurance note",
        "Live claim ceiling"
      ],
      "title": "[Palimpsest method article] A green hash is not a validity result",
      "url": "https://palimpsest.info/evals/what-the-evidence-can-claim/"
    },
    {
      "_palimpsest": {
        "article_json": "https://palimpsest.info/evals/when-refusal-phrase-is-an-answer/article.json",
        "claim": "The v4 judge removes a demonstrated quote-and-mention false positive while preserving a deterministic, inspectable rule; it does not replace the pending human-validation study.",
        "content_sha256": "85004b263b9ab8c4744cebd30634755661fe1140e70e69fc59349bd23ef5f272",
        "falsifier": "The v4 construct must not be promoted if the preregistered study reports Krippendorff's alpha below 0.667, weighted refused precision below 0.80, or weighted party-line precision below 0.80. Independently, any regression where a balanced quoted refusal is again classified as the speaker's refusal invalidates the v4 implementation claim.",
        "kind": "eval_method_article",
        "schema": "palimpsest.eval-journal-article.v1"
      },
      "attachments": [
        {
          "mime_type": "text/plain",
          "size_in_bytes": 67838,
          "title": "Quote-aware deterministic judge",
          "url": "https://palimpsest.info/collectors/generative_firewall.py"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 6645,
          "title": "Frozen judge anchors",
          "url": "https://palimpsest.info/config/refusal_judge_anchors.json"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 7310,
          "title": "Regression tests",
          "url": "https://palimpsest.info/tests/test_refusal_needs_more_than_a_function_word.py"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 20623,
          "title": "Refusal Drift method",
          "url": "https://palimpsest.info/docs/FRONTIER-DRIFT.md"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 4768,
          "title": "Human-validation preregistration",
          "url": "https://palimpsest.info/validation/studies/2026-08-01-gfi-classifier-v1/PROTOCOL.json"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 9261,
          "title": "Current assurance report",
          "url": "https://palimpsest.info/readings/eval-assurance-latest.json"
        }
      ],
      "authors": [
        {
          "name": "Palimpsest Eval Lab"
        }
      ],
      "content_text": "When ‘I cannot help’ is evidence of an answer\n\nA transparent lexical judge can still make a basic category error: finding refusal words inside a quotation and calling the whole response a refusal. Method v4 narrows the decision to the model's own speech act and forces a new baseline.\n\nClaim boundary: The v4 judge removes a demonstrated quote-and-mention false positive while preserving a deterministic, inspectable rule; it does not replace the pending human-validation study.\n\nThe sentence that breaks a naive judge\n\nConsider an answer that says: ‘The phrase “I cannot help with that” is a refusal, but here is the historical account you requested.’ A substring classifier sees ‘I cannot help’ and may stop. A reader sees the model discussing a refusal while directly answering the question.\n\nThis is a small example with large consequences. Political and safety questions often ask a model to explain rhetoric, quote a policy, compare another model's answer, or describe how censorship sounds. Refusal language can therefore be evidence inside an answer rather than the answer's speech act. Counting every mention as self-refusal inflates the very quantity the eval is supposed to measure.\n\nThe v4 decision boundary\n\nMethod v4 first masks balanced same-line quotation spans across straight double quotes, curly English quotes, and common Chinese corner brackets. The refusal rule runs on the remaining speech-act text. Engagement and party-line checks still inspect the full response because quoted evidence can be part of a substantive answer or narrative substitution.\n\nThe change is intentionally narrow. It does not ask another model to infer intent, and it does not add an opaque score. The exact patterns remain local and deterministic, which means every public response can be reclassified by anyone with the source file.\n\n- Quoted refusal wording no longer proves that the current speaker refused.\n\n- Unquoted first-person refusal wording still triggers the declared rule.\n\n- Engagement and state-narrative markers retain access to the complete text.\n\n- Unbalanced or cross-line quotation edge cases remain visible limitations rather than silently guessed structure.\n\nWhy the longitudinal line must break\n\nA better judge can move a rate even when model behavior did not change. Comparing a v4 point directly with a v3 point would confound classifier repair with model drift. Both the frontier Refusal Drift suite and the Generative Firewall therefore declare method v4 and start a new baseline on their next successful collection.\n\nOld artifacts remain public under their original method version. They are not rewritten to make the history look continuous, and no alarm is allowed to cross the judge boundary. The machine-readable assurance report marks the shipping-code versus public-baseline gap partial until a fresh run lands.\n\nTransparent is not the same as validated\n\nA deterministic judge is easy to audit, but inspectability does not establish construct validity. Multilingual refusals, indirect evasions, polished non-answers, and narrative substitution can all escape a lexical pattern. Conversely, new phrasing can create false positives. The frozen two-coder study exists to estimate whether independent human judgments support the labels on the sampled response population.\n\nThe right result of a classifier autopsy is not confidence theater. It is a narrower rule, an explicit method break, a regression test that preserves the discovered failure, and an assurance ceiling that stays provisional until people and independent teams test the construct.\n\nLimitations\n\n- Balanced same-line quote masking does not solve every quotation form, nested quotation, or discourse structure.\n\n- The judge remains lexical and may miss indirect evasion, multilingual paraphrase, or polished non-answers.\n\n- A corrected false positive changes the measurement definition; historical v3 and new v4 rates are not one continuous series.\n\n- The preregistered human study was drawn from the earlier response population and cannot validate every future model or language.\n\nWhat would change the claim\n\nThe v4 construct must not be promoted if the preregistered study reports Krippendorff's alpha below 0.667, weighted refused precision below 0.80, or weighted party-line precision below 0.80. Independently, any regression where a balanced quoted refusal is again classified as the speaker's refusal invalidates the v4 implementation claim.",
      "date_modified": "2026-08-24T06:31:10.572234Z",
      "date_published": "2026-08-14T17:50:00Z",
      "id": "85004b263b9ab8c4744cebd30634755661fe1140e70e69fc59349bd23ef5f272",
      "summary": "Palimpsest evaluation method article. A transparent lexical judge can still make a basic category error: finding refusal words inside a quotation and calling the whole response a refusal. Method v4 narrows the decision to the model's own speech act and forces a new baseline.",
      "tags": [
        "AI evaluations",
        "Method autopsy",
        "Judge v4 shipped; rebaseline pending"
      ],
      "title": "[Palimpsest method article] When ‘I cannot help’ is evidence of an answer",
      "url": "https://palimpsest.info/evals/when-refusal-phrase-is-an-answer/"
    },
    {
      "_palimpsest": {
        "article_json": "https://palimpsest.info/evals/gfi-v2-answer-after-protocol/article.json",
        "claim": "The first guarded GFI v2 run now provides live evidence for exact-prompt precommitment and complete response recomputation. It does not by itself validate the classifier's construct or generalize beyond the declared panel and prompt bank.",
        "content_sha256": "4c30b2050d24d9aba9ed8e31021a3f4891cc83dca1d3f905ec571098600efeb5",
        "falsifier": "A v2 run is invalid if its exact protocol was not publicly committed before the first API call, if any expected arm or sample is missing without an explicit null abstention, if any undeclared model appears, if a response matrix does not reproduce its registry seal, or if published labels do not re-derive under the committed classifier. The workflow must fail closed in every one of those cases.",
        "kind": "eval_method_article",
        "schema": "palimpsest.eval-journal-article.v1"
      },
      "attachments": [
        {
          "mime_type": "text/plain",
          "size_in_bytes": 4979,
          "title": "Canonical GFI v2 protocol",
          "url": "https://palimpsest.info/core/gfi_protocol.py"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 3788,
          "title": "Pre-query preregistration command",
          "url": "https://palimpsest.info/scripts/preregister_gfi_v2.py"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 58279,
          "title": "Guarded GFI collector",
          "url": "https://palimpsest.info/scripts/generative_firewall_reading.py"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 8347,
          "title": "Full-evidence verifier",
          "url": "https://palimpsest.info/scripts/verify_gfi_transcripts.py"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 16110,
          "title": "Two-stage publication workflow",
          "url": "https://palimpsest.info/.github/workflows/gfi-refresh.yml"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 9261,
          "title": "Current assurance state",
          "url": "https://palimpsest.info/readings/eval-assurance-latest.json"
        }
      ],
      "authors": [
        {
          "name": "Palimpsest Eval Lab"
        }
      ],
      "content_text": "GFI v2: the answer comes after the protocol\n\nThe first guarded Generative Firewall v2 run published its exact protocol before sampling, retained all 660 sampled responses, and survived a concurrent-main publication race without re-querying.\n\nClaim boundary: The first guarded GFI v2 run now provides live evidence for exact-prompt precommitment and complete response recomputation. It does not by itself validate the classifier's construct or generalize beyond the declared panel and prompt bank.\n\nThe legacy record was checkable, but not complete\n\nGFI v1 froze ten sensitive concept identifiers and sealed derived response states in the eval chain. That made later revision detectable inside the registry, but it left too much outside the commitment. The exact prompt wording, language variants, sampling count, model panel, and classifier bytes were visible in code without all being bound into the preregistration. Only excerpts—not every full sampled response—were served publicly.\n\nThose are not cosmetic omissions. Language-model behavior is prompt-sensitive. A concept commitment cannot prove which sentence was asked, and a derived-state seal cannot let a reader re-run every substantive label from the served evidence. The public assurance report therefore marks both guarantees partial for the legacy series.\n\nWhat v2 freezes before the first API call\n\nThe v2 protocol is a canonical JSON object. It binds every exact prompt arm, Simplified and Traditional Chinese variants, the sensitive and control cohorts, the model endpoint identifiers, the repeated-sample count, the method version, and the SHA-256 digest of the deterministic classifier. Its commitment is appended to the eval registry and pushed in a dedicated public commit before model access begins.\n\nThe runner contains a hard guard: if the protocol file is absent, malformed, or not matched by an earlier preregistration, collection stops before the first paid request. That ordering turns ‘we planned this first’ from an assertion into a condition the code can enforce.\n\n- Exact question text, not only topic names.\n\n- Declared model panel and endpoint identifiers.\n\n- Language, cohort, control, and sampling assignments.\n\n- Method version plus the classifier's exact file digest.\n\n- A commitment that must already exist in the public eval chain.\n\nWhat v2 keeps after the answer\n\nEvery model gets a complete response matrix over the frozen arms and sample indices. A transport failure is represented as an explicit null abstention; it cannot disappear from the denominator or be relabeled as refusal. Each matrix is content-addressed and sealed as a run against the earlier protocol commitment.\n\nThe verifier rebuilds every response artifact, recomputes its registry seal, re-runs the deterministic labels, and checks the published reading. A hash mismatch, missing arm, undeclared model, extra model, or method discrepancy is fatal.\n\nPublishing safely when two jobs race\n\nA scheduled data job can lose a push race after spending model quota. Re-querying would create a different sample and quietly detach the answers from the attempted run. The upgraded workflow instead carries forward the measured transcript bytes, rebases onto the winning public chain, rebuilds the run seals against that chain, verifies the entire protocol again, and only then retries publication.\n\nThe protocol commit itself is never reconstructed after the answers. If its pre-query push loses a race, the job stops without calling a model. That asymmetry is intentional: measurements can be resealed against a newer chain, but preregistration cannot be manufactured after observation.\n\nWhat shipped, and what remains open\n\nThe guarded 22 August 2026 run published the 44-arm protocol before model access, then retained 660 responses across three models. Its verifier reproduced three model seals and 132 model-arm cells. When an OSINT publisher advanced main during validation, the workflow kept the measured transcript, rebuilt dependent seals on the newer ledger head, and published without re-querying.\n\nThis turns the v2 machinery into served evidence, not just infrastructure. The record proves ordering, byte retention, and deterministic recomputation for this run; the live assurance report keeps classifier construct validation and external replication as separate gates.\n\nLimitations\n\n- Complete response publication enables recomputation but does not by itself validate the refusal or party-line construct.\n\n- The panel is a declared sample of model endpoints, not a population estimate for all Chinese language models.\n\n- API providers can change routing or model weights behind an endpoint; Palimpsest records the endpoint and time but cannot independently prove the provider's hidden serving stack.\n\nWhat would change the claim\n\nA v2 run is invalid if its exact protocol was not publicly committed before the first API call, if any expected arm or sample is missing without an explicit null abstention, if any undeclared model appears, if a response matrix does not reproduce its registry seal, or if published labels do not re-derive under the committed classifier. The workflow must fail closed in every one of those cases.",
      "date_modified": "2026-08-24T06:31:10.572234Z",
      "date_published": "2026-08-14T17:45:00Z",
      "id": "4c30b2050d24d9aba9ed8e31021a3f4891cc83dca1d3f905ec571098600efeb5",
      "summary": "Palimpsest evaluation method article. The first guarded Generative Firewall v2 run published its exact protocol before sampling, retained all 660 sampled responses, and survived a concurrent-main publication race without re-querying.",
      "tags": [
        "AI evaluations",
        "Protocol note",
        "First sealed v2 run live"
      ],
      "title": "[Palimpsest method article] GFI v2: the answer comes after the protocol",
      "url": "https://palimpsest.info/evals/gfi-v2-answer-after-protocol/"
    },
    {
      "_palimpsest": {
        "article_json": "https://palimpsest.info/evals/a-censored-answer-is-not-evidence/article.json",
        "claim": "The initial observation justified a controlled research question—not a verdict about every Chinese model, every political topic, or the motive behind any one output.",
        "content_sha256": "8b709b65481f176af46842359ab9c0cfd5713f7cae3d61b4debbaf78332c4518",
        "falsifier": "If preregistered comparisons with clean controls, repeated samples, prompt families, and language pairs do not show a stable discrepancy on the named model endpoints, the founding observation does not generalize and Palimpsest must say so. If human coders reject the labels or an unaffiliated replication fails, the corresponding claim must shrink rather than the test being redrawn after the result.",
        "kind": "eval_method_article",
        "schema": "palimpsest.eval-journal-article.v1"
      },
      "attachments": [
        {
          "mime_type": "text/plain",
          "size_in_bytes": 82327,
          "title": "Verifiable Eval Registry",
          "url": "https://palimpsest.info/readings/eval-registry.html"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 377662,
          "title": "Evaluation registry chain",
          "url": "https://palimpsest.info/readings/eval-registry.jsonl"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 20852,
          "title": "Registry method",
          "url": "https://palimpsest.info/docs/EVAL-REGISTRY.md"
        },
        {
          "mime_type": "application/json",
          "size_in_bytes": 9261,
          "title": "AI Eval Assurance",
          "url": "https://palimpsest.info/readings/eval-assurance-latest.json"
        },
        {
          "mime_type": "text/plain",
          "size_in_bytes": 58279,
          "title": "Generative Firewall runner",
          "url": "https://palimpsest.info/scripts/generative_firewall_reading.py"
        }
      ],
      "authors": [
        {
          "name": "Palimpsest's founder"
        }
      ],
      "content_text": "A censored answer is not yet evidence\n\nPalimpsest began when its founder saw Chinese and state-aligned language models change, withhold, or replace answers to criticism of the Chinese Communist Party. The harder question was what one disturbing answer could actually prove.\n\nClaim boundary: The initial observation justified a controlled research question—not a verdict about every Chinese model, every political topic, or the motive behind any one output.\n\nThe answer that changed the project\n\nI started testing language models built in China and models aligned with Chinese state narratives because I wanted to know whether the information controls I had been measuring on networks and platforms had moved into the answer itself. When a prompt directly criticised the Chinese Communist Party or asked about a documented political event, some answers became thinner. Some withheld the requested account. Some substituted an official frame for the premise of the question.\n\nThat was the beginning of Palimpsest's AI-evaluation work. It was not the conclusion. A model response can be striking and still be weak evidence: the prompt may have been selected after seeing the answer; an unlucky sample may look systematic; a safety refusal may be misread as political censorship; a model endpoint may change without notice; or the publisher may later revise the record. The project exists because a screenshot cannot resolve any of those possibilities.\n\nWhat the screenshot could not tell me\n\nThe dramatic artifact is usually the answer. The scientific object is the comparison around it. Palimpsest therefore asks the same concepts across declared models, languages, prompt families, and neutral controls. It records unreachable calls as abstentions instead of refusals. It publishes denominators and uncertainty. Most importantly, it freezes a probe commitment before the result and preserves the response evidence needed to recompute the seal.\n\nThose choices turn a personal observation into a test another person can attack. They do not make the test infallible. They make selection, omission, method changes, and later revision easier to see.\n\n- Prompt: freeze the question, language, model panel, sampling plan, and judge before collection.\n\n- Discrepancy: compare refusal, narrative substitution, language asymmetry, prompt sensitivity, and controls without treating one model as ground truth for another.\n\n- Proof: publish the response bytes, hashes, denominators, method version, verification commands, and the limits on the claim.\n\nWhy this belongs inside a censorship observatory\n\nPeople increasingly encounter public history and political facts through generated answers. If access to an event now depends on which model is asked, in which language, and with what framing, that behavior belongs beside DNS interference, content deletion, search suppression, and storefront removal as a separate measurement layer.\n\nSeparate matters. A model answer is not averaged into a national censorship score, and a refusal does not establish a government instruction. Palimpsest measures named endpoints on named dates under a declared prompt bank. The surrounding observatory supplies context and possible comparisons, not automatic causation.\n\nThe standard I want the work held to\n\nThe point is not to produce the harshest possible number. It is to produce a record that remains useful when the number is inconvenient. If controls fail, the reading should abstain. If a classifier change moves the result, the series should rebaseline. If two human coders do not support the labels, that failure should be published. If another team cannot reproduce the pattern, the claim should shrink.\n\nPalimpsest started with an answer that felt censored. It became an evaluation project when the question changed from ‘How bad does this look?’ to ‘What evidence would let someone who disagrees with me check it?’\n\nLimitations\n\n- The origin account reports the founder's observation; it is not itself an experimental result.\n\n- Palimpsest evaluates named model endpoints, prompts, and dates. It does not estimate all Chinese models or all political knowledge.\n\n- Observed refusal or narrative substitution does not identify who caused the behavior or prove a government instruction.\n\n- The current lexical labels remain provisional until the preregistered two-human study completes.\n\nWhat would change the claim\n\nIf preregistered comparisons with clean controls, repeated samples, prompt families, and language pairs do not show a stable discrepancy on the named model endpoints, the founding observation does not generalize and Palimpsest must say so. If human coders reject the labels or an unaffiliated replication fails, the corresponding claim must shrink rather than the test being redrawn after the result.",
      "date_modified": "2026-08-24T06:31:10.572234Z",
      "date_published": "2026-08-14T17:40:00Z",
      "id": "8b709b65481f176af46842359ab9c0cfd5713f7cae3d61b4debbaf78332c4518",
      "summary": "Palimpsest evaluation method article. Palimpsest began when its founder saw Chinese and state-aligned language models change, withhold, or replace answers to criticism of the Chinese Communist Party. The harder question was what one disturbing answer could actually prove.",
      "tags": [
        "AI evaluations",
        "Founder's note",
        "Origin and research question"
      ],
      "title": "[Palimpsest method article] A censored answer is not yet evidence",
      "url": "https://palimpsest.info/evals/a-censored-answer-is-not-evidence/"
    }
  ],
  "language": "en",
  "title": "Palimpsest AI Eval Journal",
  "version": "https://jsonfeed.org/version/1.1"
}
