Skip to content

For researchers, journalists and analysts

China internet censorship data and AI evaluation evidence

Palimpsest is a free public-good observatory for researching the Great Firewall, documented deletion and AI-model refusal behavior. It keeps network, content and model evidence separate, publishes source files and limitations beside every reading, and watches the censor—never the censored.

For journalists: find, trail, export, cite

A reporter who is not the operator should be able to find a public deletion, see the evidence trail, download it, and cite it. The erasure desk is that path. It is not buried in a JSON blob.

Find a deleted post

Open the journalist evidence trail. Each row is an already-public post, ledger item, or Wayback reconstruction. Search the table by title, term, or source URL. Absence is a coverage gap, not a claim that nothing was deleted.

See the trail

Every row shows first-seen, last-seen, last-confirmed-alive when known, the source URL, a Wayback snapshot or lookup, and a SHA-256 of the public excerpt. The Situation desk repeats those fields on linked OSINT rows. A lookup URL is an address to try, not a claimed capture, unless a snapshot is attached.

Export and cite

Download the CSV or JSON. Copy the citation line on the row. Cite Palimpsest as the observatory that recorded a public disappearance — not as a witness inside China and not as proof of motive. Method: journalist guide.

What is captured, and what is not

Captured: public posts, public deletion ledgers, Wayback reconstructions, and GFW injector telemetry from separate instruments. Not captured: private WeChat, classified systems, in-country accounts, follower graphs, comments, locations, or media binaries. Palimpsest never fabricates a live reading when a collector is silent.

Plain answers for researchers

What data can measure the Great Firewall of China?

Use more than one vantage. Palimpsest publishes separate measurements from OONI and Censored Planet and from volunteer DNS probes. When methods disagree, the result is an interval—not a fabricated national rate.

How can internet deletion and censorship be measured?

Documented removals can be treated as observations, then compared over time for changing attention and novelty. The DDTI Observatory exposes the samples behind each reading. A deletion does not, by itself, prove motive or measure every act of censorship.

What does an AI model refusal rate prove?

It describes behavior on a frozen prompt suite at a stated time. It does not prove motive or universal behavior. The Eval Registry seals prompts, outputs and labels so others can verify the record and avoid comparing incompatible suites; the assurance report states which stronger validity claims remain open.

Live signals

Each signal recomputes on a fixed schedule and publishes a machine readable file plus a committed time series. Signals that would need a network vantage inside China are held back on purpose until that measurement can be verified, rather than published as a guess.

SignalWhat it measuresSourceUpdatedStatus
DDTI
deletion-differential threat index
Which topics the censor is most actively scrubbing, ranked by attention and novelty. China Digital Times every 3 hours LIVE
Generative Firewall
state-AI refusal index
The share of sensitive prompts that state-aligned LLMs refuse or rewrite, against neutral controls. Public model APIs daily LIVE
GDELT cross-signal
censored at home, loud abroad
Which censored topics the world's press is covering heavily (containment) versus silently absent (blackout). GDELT global news index every 6 hours LIVE
GitHub-refuge
pressure on the mirrors
Takedowns, legal blocks, and visibility drops on GitHub repos that shelter censored material, plus defensive fork and star bursts. GitHub public API twice a day LIVE
Eval Registry
sealed model evaluations
Pre-registered, hash-chained audits of Chinese state-aligned and Western frontier models, including refusal drift over time. The registry keeps separate suites: the preserved cn-sensitive-generative-firewall-v1 history covers the state-aligned panel, its staged v2 protocol adds exact-prompt and full-matrix evidence, and frontier-overrefusal-v2 covers the Western frontier panel. They share evidence machinery, not a rubric. A served edit breaks the chain; public history, anchors and witnesses address whole-chain rewriting. The v2 frontier suite asks questions through paraphrase families, publishes Wilson intervals, and seals raw-response digests so a reader can reproduce the current seals and labels. The assurance report separately shows that the human study and unaffiliated replication are unfinished. Public model APIs every 6 hours LIVE
Erasure Observatory
composite erasure index
What the record lost, across the network, narrative, and model layers, sealed into a tamper-evident ledger. Layers that cannot report are shown absent, never zero-filled. OONI and model audits; Baike disabled and retained as stale every 6 hours LIVE
GFW network signal
live firewall blocking
Independent, side-channel measurements of Great Firewall blocking events, with historical backfill. OONI open data every 6 hours LIVE
Velocity
deletion speed
How fast a post is deleted after posting, timed to the minute. In-country observation held back SUPPRESSED

Get the data

All files are plain JSON. The *-latest.json files hold the current snapshot; the *-history.jsonl files are append-only time series, one compact record per run, so you can chart a signal over time.

osint-china-latest.json
The complete OSINT China bundle: every China-facing current payload, normalized source health, cadence, freshness deadline, layer coverage and honest missing-source state in one versioned contract.
ddti-latest.json
DDTI snapshot: ranked terms with threat, attention, novelty, is_new, and source samples with links.
ddti-history.jsonl
DDTI time series: generated_at, n_terms, top_term, top_threat per run.
latest.json
Generative Firewall raw run: per-model, per-concept refusals, controls, and the GFI value.
history.jsonl
Generative Firewall time series: date, gfi, censored, total, controls.
gdelt-latest.json
GDELT cross-signal: each DDTI term with global coverage volume and a containment or blackout label.
gdelt-history.jsonl
GDELT time series: n_terms, containment and blackout counts, top term.
github-refuge-latest.json
GitHub-refuge: each watched mirror with status, pressure likelihood, fork and star novelty, and any matched takedown notices.
github-refuge-history.jsonl
GitHub-refuge time series: repos watched, present, and pressure events per run.
eval-registry.jsonl
The registry chain itself: every pre-registration and sealed run, hash-chained in order. This file is the tamper-evident record; verify it with scripts/verify_eval_registry.py.
eval-registry-latest.json
Registry summary: attestation counts, models audited, recent runs, Merkle root, head hash, verified flag.
eval-assurance-latest.json
Claim-by-claim assurance: integrity, prompt precommitment, raw-response recomputation, pipeline reproducibility, statistics, human construct validation and independent replication. It publishes a claim ceiling rather than blending unlike guarantees into one score.
refusal-drift-latest.json
Frontier refusal drift: per-model refusal rate over question families with a Wilson 95% interval, wording invariance across three phrasings per family, matched English/Chinese asymmetry, drift versus the prior comparable run, and the anytime-valid churn alarm. Method and honest limits in docs/FRONTIER-DRIFT.md.
refusal-drift-transcripts.json
The raw model responses behind the current refusal-drift reading, with the prompt that drew each. Their digests are what the sealed run commits to, so scripts/verify_refusal_transcripts.py can prove the served text is the sealed text and re-derive every label. Current run only; prior runs live in git history and still verify against their seals.
refusal-drift-churn.jsonl
One line per model per comparable run: how many probes changed label and how many were compared. This is the evidence the standing alarm accumulates, kept separate from the findings history because a null calibrated only on runs where something moved would be estimated from a sample selected on having moved.
erasure-observatory-latest.json
Composite Erasure Index snapshot: per-layer readings (network, narrative, model) and the mean of reporting layers.
erasure-ledger.jsonl
The sealed erasure ledger: every observatory reading, hash-chained; verify with scripts/verify_ledger.py.
ooni-gfw-latest.json
GFW network signal: independent side-channel blocking measurements from OONI open data.
board-alarm-latest.json
What the board as a whole concludes, with the multiplicity paid for: e-BH selection across every monitored signal (false-discovery control under arbitrary dependence), a board-wide merged e-value, and how many of the network/content/model layers are elevated at once. A stack of per-signal guarantees is not a board guarantee; this file is the board one. The model layer runs on the v2 refusal suite's own anytime-valid churn e-process (the closed v1 conformal series is listed for the record and contributes no evidence).
coverage-guard-latest.json
Whether each signal's movement survives conditioning on its own sample size. A censorship rate can fall because probes thinned out rather than because censorship eased; verdicts are CONFIRMED, COVERAGE_CONFOUNDED or NO_MOVE, each with the fit and the raw-vs-conditional effect size. Check this before quoting any signal as a change.
forecast-ledger-latest.json
Our own scored track record. Every signal forecast one step ahead from only its past (strictly prequential, never refit with hindsight), scored by the Weighted Interval Score against what actually arrived. Includes empirical coverage against nominal, skill against a baseline, and the worst misses by date — the misses are not removable.
cross-layer-latest.json
Does one layer of the apparatus move before another? Lead/lag between network, content and model signals against a circular-shift null on differenced series, Holm-corrected. Reports timing, never cause. Publishes the false-positive rates of the naive test (98%) and of this one (8%) so the choice of null is auditable.
event-flags-latest.json
Per-signal anytime-valid change alarms: two-sided conformal Shiryaev-Roberts e-detectors, so a signal collapsing flags as well as one rising, with direction and effect size on every reading.
vantage-fusion-latest.json
OONI and Censored Planet fused into one GFW reading as an INTERVAL, not a point. When the two methods diverge the range widens and single_rate_quotable goes false, because a midpoint of two numbers that disagree is not an estimate.

A machine-readable summary for AI agents lives at llms.txt. The full directory is at /readings.

How it is measured

Three layers, one instrument Network, content and model layers each feed their own collectors. Collectors publish readings as JSON with a history file beside each. The board computes alarms, intervals and forecasts from the readings, and every reading and board state is sealed into a hash-chained ledger whose roots are anchored outside the project. network layer probes · OONI · storefronts content layer deletions · blocklists model layer frozen suites · sealed runs collectors scheduled · gated published readings JSON + history files the board alarms · intervals sealed ledger chained · anchored
Each layer has its own vantage and its own failure modes, so each is collected, published and cross-checked on its own terms — and everything downstream of the readings can be recomputed by anyone from the published files alone.

Our own misses

Every signal on this board is also forecast one step ahead from only its own past, never refit with hindsight, and then scored against what actually arrived. The uncomfortable half is the point: the published interval either covered the next reading or it did not, and both counts are here. A forecast record with the misses taken out is worth nothing.

The scored track record is read from forecast-ledger-latest.json when this page loads. Figures appear below once that file is read. If it cannot be read, this panel says so instead of showing a number.

Scope and honest limits

Palimpsest publishes only what it can stand behind. It measures censor attention, which topics draw the most suppression effort, not a full deletion rate. The velocity signal, how fast a specific post is removed, needs a network vantage inside China to time deletions to the minute. Until that vantage is in place and verified, velocity stays suppressed by design. Nothing on this site is a modelled guess dressed up as a measurement.

How to cite

If you use this data in an article, report, or paper, a citation and a link back are appreciated. Please cite the accessed date, since the signals update continuously. For a specific signal on a specific day, use the citation builder or python3 -m scripts.build_citation_pack --dataset ddti --day YYYY-MM-DD. To challenge a number, follow How to challenge a number. The sealed weekly fusion lives at weekly-situation.html.

Palimpsest (2026). Palimpsest Censorship Observatory and Verifiable Eval
Registry: DDTI, Generative Firewall Index, GDELT cross-signal, and sealed
model evaluations [live dataset]. https://palimpsest.info
(accessed YYYY-MM-DD).
@misc{palimpsest,
  title  = {Palimpsest Censorship Observatory and Verifiable Eval Registry},
  author = {Palimpsest},
  year   = {2026},
  url    = {https://palimpsest.info},
  note   = {Open live dataset: DDTI, Generative Firewall Index, GDELT
            cross-signal, Wayback deletion reconstruction, sealed eval
            registry with refusal drift}
}

Reproducibility and provenance

Palimpsest is fully open source under the MIT licence. The signals recompute inside public, auditable GitHub Actions, and every refresh is a timestamped commit in the repository, so the entire history of what was measured, and when, is public and verifiable. There is no private backend deciding the numbers.

Two of the surfaces go further than the commit trail: the erasure ledger and the eval registry are hash-chained and Merkle-committed, so they stay verifiable even outside git. Clone the repo and run python3 scripts/verify_ledger.py or python3 scripts/verify_eval_registry.py; exit 0 means every seal recomputes and, for the registry, that every run referenced a probe set frozen before the model was queried. This constraint applies to us too. If we edited a published number, our own verifier would report the break.

The roots are also deposited with parties we do not control. Every refresh that moves a root gets an Internet Archive snapshot of the served chain files and an OpenTimestamps stamp into Bitcoin (the .ots proofs live in readings/anchors/ and verify with the standard client, against the blockchain, not against us; anchors-latest.json is the currently stamped root). An independent witness on separate infrastructure re-verifies the served chains on a timer and alerts if any previously seen history changes; anyone can run one with python3 ops/witness/palimpsest_witness.py. Single attestations verify without downloading the chain via python3 scripts/prove_inclusion.py <seq>. The full layer-by-layer trust model, including what these layers cannot prove, is written down in docs/INTEGRITY.md.

Read Chinese? We need four hours of your judgment

The Generative Firewall Index labels every model answer with a transparent rule based classifier: refused, state narrative, or answered. A researcher whose work this instrument builds on asked us a fair question: would actual humans agree with those labels? We want to find out properly, and that part cannot be automated. We need two volunteer coders.

The task. We send you a short manual and a spreadsheet of 145 model answers, some Chinese, some English. Each one is a chatbot's answer to a question. Lengths vary a lot: over two fifths of them are under 500 characters, the median is 683, and the longest is 3,509, so a few are a screen or two of reading. You read the answer and pick one of three labels. That's the whole job. Around four hours, alone, on your own schedule inside two weeks.

Why 145 and not more. The sampler asked for 60 refusals and 50 party-line answers and could only draw 17 and 38: outright refusals and undisguised state-narrative answers are simply rare in the corpus the models produced. We shipped the shortfall rather than pad those cells with near misses, so agreement on the refused and party_line strata will carry wide uncertainty and we will report it that way. The exact targets, pool sizes, and shortfalls are recorded in validation/studies/2026-08-01-gfi-classifier-v1/manifest.json.

Who. Anyone who reads Chinese fluently. Any nationality, no technical background, the manual carries everything you need. One hard restriction: if you are currently in mainland China or Hong Kong, please do not volunteer. The texts touch politically sensitive topics, and no dataset is worth risk to you.

The rules. Work alone, no comparing notes with the other coder until you are both done, and no AI help of any kind. If a machine assists the judgment, the study measures nothing. The entire point is unaided human agreement.

What you get. Your name in the acknowledgments here and in the research paper this feeds, or full anonymity if you prefer. Either way, a public instrument that watches censorship gets a stronger spine because you read carefully for an afternoon.

To volunteer, open a GitHub issue titled Validation coder and say roughly where you are and how you come by your Chinese: native, degree, HSK, or grew up with it. First two qualified volunteers get the sheets.

Work with us

If you are a reporter, researcher, or think tank and want a specific term tracked, a data extract, or help interpreting a signal, open an issue on GitHub. Palimpsest is a public good and collaboration is welcome.

Free and open source, developed in the open as a public good. Never a commercial product, and it never monetizes the people or topics it observes. The internet stays free because people keep measuring the dark.