<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel>
  <title>Palimpsest AI Eval Journal</title>
  <link>https://palimpsest.info/evals/</link>
  <description>Evidence-bound essays about censorship evaluations, method changes, failures, and what the current record can actually support.</description>
  <language>en</language>
  <lastBuildDate>Mon, 24 Aug 2026 06:31:10 +0000</lastBuildDate>
  <atom:link href="https://palimpsest.info/evals/feed.xml" rel="self" type="application/rss+xml" />
  <item>
    <guid isPermaLink="false">urn:sha256:481deb73f5ae136a08961ce138dd544cfdde949f83a0e79943db0c80b202c8af</guid>
    <title>[Palimpsest method article] A green hash is not a validity result</title>
    <link>https://palimpsest.info/evals/what-the-evidence-can-claim/</link>
    <pubDate>Fri, 14 Aug 2026 17:54:00 +0000</pubDate>
    <description>Palimpsest evaluation method article. Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.</description>
    <category>palimpsest-eval-method</category>
    <content:encoded><![CDATA[A green hash is not a validity result

Palimpsest now publishes a machine-readable assurance ladder that keeps chain integrity, exact-prompt commitment, response recomputation, statistical design, human validation, and independent replication on separate axes.

Claim boundary: The current eval record supports a provisional measurement claim; it does not yet support calling the lexical construct human-validated or the findings independently replicated.

The most dangerous badge was the easiest one to earn

Hash chains answer a valuable, narrow question: do the served records reproduce their commitments and publication order? They do not answer whether a prompt measures the intended construct, whether a classifier agrees with people, whether a panel represents a larger model population, or whether another team can reproduce the finding.

Collapsing those questions into one ‘verified’ badge rewards the easiest engineering property and launders the hardest scientific gaps. Palimpsest's assurance report refuses that composite. It exposes seven dimensions and assigns each check a status, evidence statement, limitation, affected suite, and—where possible—a local verification command.

Seven questions, no average

Integrity asks whether the registry recomputes and whether every run follows a preregistration. Prompt precommitment asks whether the exact questions were fixed before answers. Response recomputability asks whether full public text can reproduce the current seals. Pipeline reproducibility asks whether labels can be regenerated under an identified method. Statistical design asks for denominators, uncertainty, controls, power, and correction for repeated looks. Construct validation asks whether independent humans support what the labels mean. Replication asks whether an unaffiliated team reproduced the result.

A failure in one dimension is not averaged away by passes elsewhere. The dimension inherits its weakest applicable check, and the public claim ceiling follows the unresolved scientific gates.

- Pass: the served evidence satisfies the declared machine check.

- Partial: a real guarantee exists, but a named part of it is absent or belongs to a legacy method.

- Pending: a preregistered study or evidence-producing step is not complete.

- Open: the invitation exists, but no qualifying independent result is on record.

- Fail: published evidence exists and violates the declared contract or frozen threshold.

The human-validation gate cannot be faked by a filename

The two-coder study is already frozen at 145 rows with a codebook, sample commitment, weighting plan, and three rejection thresholds. Assurance stays pending until one exact result artifact identifies that commitment, accounts for every row from both coders, carries both attestations, and supplies digests that match the released coder sheets, answer key, manifest, and protocol.

Once complete, the check independently recomputes whether Krippendorff's alpha reaches 0.667 and whether weighted precision reaches 0.80 for both refused and party-line labels. A malformed result is fail, not pending. A threshold miss is a published failure, not permission to redraw the sample.

What can be said today

Today Palimpsest can say that its public eval chain is tamper-evident under the declared hash construction; the frontier suite binds exact prompts and exposes full current responses for seal recomputation; and both suites publish explicit denominators, uncertainty, controls, and method versions. The live JSON beside this article reports the exact current count of pass, partial, pending, open, and fail checks.

It cannot yet say that the lexical classifier has completed independent human validation. The legacy China-focused GFI still lacks exact-prompt preregistration and the complete served response matrix that v2 will require. No unaffiliated team has published a preregistered replication against the registry. Those gaps are not footnotes to a strong score. They are the ceiling on the claim.

Why this makes a stronger grant case

A grant should fund a specific reduction in uncertainty, not reward a polished claim that cannot fail. The assurance ladder converts open work into auditable milestones: complete the frozen human study; land the first full-evidence GFI v2 run; recruit and support an unaffiliated replication; extend prompts and coder validation across languages; deposit durable external witnesses; and publish every miss.

That does not guarantee an award. It gives a reviewer something better than ambition: a public baseline, a named gap, a falsifier, a deliverable that changes the claim ceiling only when the evidence earns it, and a verification path that remains after the grant ends.

Limitations

- The assurance report evaluates Palimpsest's declared contracts; it is not an accreditation or an external audit.

- A pass establishes the named machine check only and must be read with that check's limitation.

- Human validation remains pending and independent replication remains open in the current public state.

- The claim ceiling can regress if later evidence fails; promotion is not permanent.

What would change the claim

Any broken registry link, prompt commitment mismatch, unrecomputable response seal, label disagreement under the committed method, missing statistical field, malformed human-study result, frozen-threshold miss, or failed qualifying replication must appear as a non-pass state and lower the affected dimension. If the public page and machine-readable report disagree, the machine-readable check is authoritative and publication should fail.]]></content:encoded>
  </item>
  <item>
    <guid isPermaLink="false">urn:sha256:85004b263b9ab8c4744cebd30634755661fe1140e70e69fc59349bd23ef5f272</guid>
    <title>[Palimpsest method article] When ‘I cannot help’ is evidence of an answer</title>
    <link>https://palimpsest.info/evals/when-refusal-phrase-is-an-answer/</link>
    <pubDate>Fri, 14 Aug 2026 17:50:00 +0000</pubDate>
    <description>Palimpsest evaluation method article. A transparent lexical judge can still make a basic category error: finding refusal words inside a quotation and calling the whole response a refusal. Method v4 narrows the decision to the model's own speech act and forces a new baseline.</description>
    <category>palimpsest-eval-method</category>
    <content:encoded><![CDATA[When ‘I cannot help’ is evidence of an answer

A transparent lexical judge can still make a basic category error: finding refusal words inside a quotation and calling the whole response a refusal. Method v4 narrows the decision to the model's own speech act and forces a new baseline.

Claim boundary: The v4 judge removes a demonstrated quote-and-mention false positive while preserving a deterministic, inspectable rule; it does not replace the pending human-validation study.

The sentence that breaks a naive judge

Consider an answer that says: ‘The phrase “I cannot help with that” is a refusal, but here is the historical account you requested.’ A substring classifier sees ‘I cannot help’ and may stop. A reader sees the model discussing a refusal while directly answering the question.

This is a small example with large consequences. Political and safety questions often ask a model to explain rhetoric, quote a policy, compare another model's answer, or describe how censorship sounds. Refusal language can therefore be evidence inside an answer rather than the answer's speech act. Counting every mention as self-refusal inflates the very quantity the eval is supposed to measure.

The v4 decision boundary

Method v4 first masks balanced same-line quotation spans across straight double quotes, curly English quotes, and common Chinese corner brackets. The refusal rule runs on the remaining speech-act text. Engagement and party-line checks still inspect the full response because quoted evidence can be part of a substantive answer or narrative substitution.

The change is intentionally narrow. It does not ask another model to infer intent, and it does not add an opaque score. The exact patterns remain local and deterministic, which means every public response can be reclassified by anyone with the source file.

- Quoted refusal wording no longer proves that the current speaker refused.

- Unquoted first-person refusal wording still triggers the declared rule.

- Engagement and state-narrative markers retain access to the complete text.

- Unbalanced or cross-line quotation edge cases remain visible limitations rather than silently guessed structure.

Why the longitudinal line must break

A better judge can move a rate even when model behavior did not change. Comparing a v4 point directly with a v3 point would confound classifier repair with model drift. Both the frontier Refusal Drift suite and the Generative Firewall therefore declare method v4 and start a new baseline on their next successful collection.

Old artifacts remain public under their original method version. They are not rewritten to make the history look continuous, and no alarm is allowed to cross the judge boundary. The machine-readable assurance report marks the shipping-code versus public-baseline gap partial until a fresh run lands.

Transparent is not the same as validated

A deterministic judge is easy to audit, but inspectability does not establish construct validity. Multilingual refusals, indirect evasions, polished non-answers, and narrative substitution can all escape a lexical pattern. Conversely, new phrasing can create false positives. The frozen two-coder study exists to estimate whether independent human judgments support the labels on the sampled response population.

The right result of a classifier autopsy is not confidence theater. It is a narrower rule, an explicit method break, a regression test that preserves the discovered failure, and an assurance ceiling that stays provisional until people and independent teams test the construct.

Limitations

- Balanced same-line quote masking does not solve every quotation form, nested quotation, or discourse structure.

- The judge remains lexical and may miss indirect evasion, multilingual paraphrase, or polished non-answers.

- A corrected false positive changes the measurement definition; historical v3 and new v4 rates are not one continuous series.

- The preregistered human study was drawn from the earlier response population and cannot validate every future model or language.

What would change the claim

The v4 construct must not be promoted if the preregistered study reports Krippendorff's alpha below 0.667, weighted refused precision below 0.80, or weighted party-line precision below 0.80. Independently, any regression where a balanced quoted refusal is again classified as the speaker's refusal invalidates the v4 implementation claim.]]></content:encoded>
  </item>
  <item>
    <guid isPermaLink="false">urn:sha256:4c30b2050d24d9aba9ed8e31021a3f4891cc83dca1d3f905ec571098600efeb5</guid>
    <title>[Palimpsest method article] GFI v2: the answer comes after the protocol</title>
    <link>https://palimpsest.info/evals/gfi-v2-answer-after-protocol/</link>
    <pubDate>Fri, 14 Aug 2026 17:45:00 +0000</pubDate>
    <description>Palimpsest evaluation method article. The first guarded Generative Firewall v2 run published its exact protocol before sampling, retained all 660 sampled responses, and survived a concurrent-main publication race without re-querying.</description>
    <category>palimpsest-eval-method</category>
    <content:encoded><![CDATA[GFI v2: the answer comes after the protocol

The first guarded Generative Firewall v2 run published its exact protocol before sampling, retained all 660 sampled responses, and survived a concurrent-main publication race without re-querying.

Claim boundary: The first guarded GFI v2 run now provides live evidence for exact-prompt precommitment and complete response recomputation. It does not by itself validate the classifier's construct or generalize beyond the declared panel and prompt bank.

The legacy record was checkable, but not complete

GFI v1 froze ten sensitive concept identifiers and sealed derived response states in the eval chain. That made later revision detectable inside the registry, but it left too much outside the commitment. The exact prompt wording, language variants, sampling count, model panel, and classifier bytes were visible in code without all being bound into the preregistration. Only excerpts—not every full sampled response—were served publicly.

Those are not cosmetic omissions. Language-model behavior is prompt-sensitive. A concept commitment cannot prove which sentence was asked, and a derived-state seal cannot let a reader re-run every substantive label from the served evidence. The public assurance report therefore marks both guarantees partial for the legacy series.

What v2 freezes before the first API call

The v2 protocol is a canonical JSON object. It binds every exact prompt arm, Simplified and Traditional Chinese variants, the sensitive and control cohorts, the model endpoint identifiers, the repeated-sample count, the method version, and the SHA-256 digest of the deterministic classifier. Its commitment is appended to the eval registry and pushed in a dedicated public commit before model access begins.

The runner contains a hard guard: if the protocol file is absent, malformed, or not matched by an earlier preregistration, collection stops before the first paid request. That ordering turns ‘we planned this first’ from an assertion into a condition the code can enforce.

- Exact question text, not only topic names.

- Declared model panel and endpoint identifiers.

- Language, cohort, control, and sampling assignments.

- Method version plus the classifier's exact file digest.

- A commitment that must already exist in the public eval chain.

What v2 keeps after the answer

Every model gets a complete response matrix over the frozen arms and sample indices. A transport failure is represented as an explicit null abstention; it cannot disappear from the denominator or be relabeled as refusal. Each matrix is content-addressed and sealed as a run against the earlier protocol commitment.

The verifier rebuilds every response artifact, recomputes its registry seal, re-runs the deterministic labels, and checks the published reading. A hash mismatch, missing arm, undeclared model, extra model, or method discrepancy is fatal.

Publishing safely when two jobs race

A scheduled data job can lose a push race after spending model quota. Re-querying would create a different sample and quietly detach the answers from the attempted run. The upgraded workflow instead carries forward the measured transcript bytes, rebases onto the winning public chain, rebuilds the run seals against that chain, verifies the entire protocol again, and only then retries publication.

The protocol commit itself is never reconstructed after the answers. If its pre-query push loses a race, the job stops without calling a model. That asymmetry is intentional: measurements can be resealed against a newer chain, but preregistration cannot be manufactured after observation.

What shipped, and what remains open

The guarded 22 August 2026 run published the 44-arm protocol before model access, then retained 660 responses across three models. Its verifier reproduced three model seals and 132 model-arm cells. When an OSINT publisher advanced main during validation, the workflow kept the measured transcript, rebuilt dependent seals on the newer ledger head, and published without re-querying.

This turns the v2 machinery into served evidence, not just infrastructure. The record proves ordering, byte retention, and deterministic recomputation for this run; the live assurance report keeps classifier construct validation and external replication as separate gates.

Limitations

- Complete response publication enables recomputation but does not by itself validate the refusal or party-line construct.

- The panel is a declared sample of model endpoints, not a population estimate for all Chinese language models.

- API providers can change routing or model weights behind an endpoint; Palimpsest records the endpoint and time but cannot independently prove the provider's hidden serving stack.

What would change the claim

A v2 run is invalid if its exact protocol was not publicly committed before the first API call, if any expected arm or sample is missing without an explicit null abstention, if any undeclared model appears, if a response matrix does not reproduce its registry seal, or if published labels do not re-derive under the committed classifier. The workflow must fail closed in every one of those cases.]]></content:encoded>
  </item>
  <item>
    <guid isPermaLink="false">urn:sha256:8b709b65481f176af46842359ab9c0cfd5713f7cae3d61b4debbaf78332c4518</guid>
    <title>[Palimpsest method article] A censored answer is not yet evidence</title>
    <link>https://palimpsest.info/evals/a-censored-answer-is-not-evidence/</link>
    <pubDate>Fri, 14 Aug 2026 17:40:00 +0000</pubDate>
    <description>Palimpsest evaluation method article. Palimpsest began when its founder saw Chinese and state-aligned language models change, withhold, or replace answers to criticism of the Chinese Communist Party. The harder question was what one disturbing answer could actually prove.</description>
    <category>palimpsest-eval-method</category>
    <content:encoded><![CDATA[A censored answer is not yet evidence

Palimpsest began when its founder saw Chinese and state-aligned language models change, withhold, or replace answers to criticism of the Chinese Communist Party. The harder question was what one disturbing answer could actually prove.

Claim boundary: The initial observation justified a controlled research question—not a verdict about every Chinese model, every political topic, or the motive behind any one output.

The answer that changed the project

I started testing language models built in China and models aligned with Chinese state narratives because I wanted to know whether the information controls I had been measuring on networks and platforms had moved into the answer itself. When a prompt directly criticised the Chinese Communist Party or asked about a documented political event, some answers became thinner. Some withheld the requested account. Some substituted an official frame for the premise of the question.

That was the beginning of Palimpsest's AI-evaluation work. It was not the conclusion. A model response can be striking and still be weak evidence: the prompt may have been selected after seeing the answer; an unlucky sample may look systematic; a safety refusal may be misread as political censorship; a model endpoint may change without notice; or the publisher may later revise the record. The project exists because a screenshot cannot resolve any of those possibilities.

What the screenshot could not tell me

The dramatic artifact is usually the answer. The scientific object is the comparison around it. Palimpsest therefore asks the same concepts across declared models, languages, prompt families, and neutral controls. It records unreachable calls as abstentions instead of refusals. It publishes denominators and uncertainty. Most importantly, it freezes a probe commitment before the result and preserves the response evidence needed to recompute the seal.

Those choices turn a personal observation into a test another person can attack. They do not make the test infallible. They make selection, omission, method changes, and later revision easier to see.

- Prompt: freeze the question, language, model panel, sampling plan, and judge before collection.

- Discrepancy: compare refusal, narrative substitution, language asymmetry, prompt sensitivity, and controls without treating one model as ground truth for another.

- Proof: publish the response bytes, hashes, denominators, method version, verification commands, and the limits on the claim.

Why this belongs inside a censorship observatory

People increasingly encounter public history and political facts through generated answers. If access to an event now depends on which model is asked, in which language, and with what framing, that behavior belongs beside DNS interference, content deletion, search suppression, and storefront removal as a separate measurement layer.

Separate matters. A model answer is not averaged into a national censorship score, and a refusal does not establish a government instruction. Palimpsest measures named endpoints on named dates under a declared prompt bank. The surrounding observatory supplies context and possible comparisons, not automatic causation.

The standard I want the work held to

The point is not to produce the harshest possible number. It is to produce a record that remains useful when the number is inconvenient. If controls fail, the reading should abstain. If a classifier change moves the result, the series should rebaseline. If two human coders do not support the labels, that failure should be published. If another team cannot reproduce the pattern, the claim should shrink.

Palimpsest started with an answer that felt censored. It became an evaluation project when the question changed from ‘How bad does this look?’ to ‘What evidence would let someone who disagrees with me check it?’

Limitations

- The origin account reports the founder's observation; it is not itself an experimental result.

- Palimpsest evaluates named model endpoints, prompts, and dates. It does not estimate all Chinese models or all political knowledge.

- Observed refusal or narrative substitution does not identify who caused the behavior or prove a government instruction.

- The current lexical labels remain provisional until the preregistered two-human study completes.

What would change the claim

If preregistered comparisons with clean controls, repeated samples, prompt families, and language pairs do not show a stable discrepancy on the named model endpoints, the founding observation does not generalize and Palimpsest must say so. If human coders reject the labels or an unaffiliated replication fails, the corresponding claim must shrink rather than the test being redrawn after the result.]]></content:encoded>
  </item>
</channel>
</rss>
