{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Palimpsest AI eval results",
  "home_page_url": "https://palimpsest.info/journal/",
  "feed_url": "https://palimpsest.info/journal/feed.json",
  "description": "Palimpsest's own dated AI evaluation results, with controls, uncertainty, deterministic authorship and revision identity attached.",
  "language": "en",
  "items": [
    {
      "id": "evalarticlev-7dff12f8c34e0a64e9910063",
      "url": "https://palimpsest.info/journal/before-reading-the-score-read-the-controls/",
      "title": "[Palimpsest eval finding] The headline was 0%. The controls still failed.",
      "summary": "Palimpsest AI evaluation finding. Meta Llama 3.3 70B Instruct refused 0 of 34 monitored question families, but 1 ordinary control prompt arms also refused. That makes the result an instrument warning, not a censorship finding.",
      "content_text": "ITEM TYPE: PALIMPSEST AI EVALUATION FINDING\n\nAuthorship: deterministic-eval-analysis\n\nA headline score is not interpretable until the same run passes its ordinary controls.\n\nTwo readings from one run\n\nOn 24 Aug 2026, Meta Llama 3.3 70B Instruct recorded 0 refused families out of 34 monitored families, a headline rate of 0%. The same run refused 1 control prompt arms across 1 ordinary control families. Its wording consistency was 97.6% across the testable families.\n\nWhy the controls outrank the score\n\nPalimpsest uses deliberately ordinary questions as controls, so a refusal there marks an instrument fault for this run. When that gate fails, the desk may describe the failure itself but may not interpret the headline rate as selective suppression.\n\nThe counterread\n\n3 of the 4 panel runs answered every control in the same sweep. That argues against calling the entire panel unusable, but it does not identify why one run failed.\n\nWhat changed since the prior comparable sweep\n\nIn the full sweep on 23 Aug 2026, the same model reported an arm refusal rate of 0.7% and controls_clean=false. Wording consistency moved from 97.6% to 97.6%. The comparison is descriptive and does not assign the change to a model release, provider, or prompt-routing decision.",
      "date_published": "2026-08-24T04:54:42.622246+00:00",
      "date_modified": "2026-08-24T04:54:42.622246+00:00",
      "authors": [
        {
          "name": "Palimpsest Eval Desk",
          "url": "https://palimpsest.info/journal/"
        }
      ],
      "tags": [
        "AI evaluation",
        "instrument-warning"
      ],
      "_palimpsest": {
        "kind": "eval_finding",
        "article_id": "evalarticle-2a851582dec97dfbf095",
        "revision_id": "evalarticlev-7dff12f8c34e0a64e9910063",
        "finding_state": "instrument-warning",
        "citation_coverage": 1.0,
        "authorship_mode": "deterministic-eval-analysis"
      }
    },
    {
      "id": "evalarticlev-a322b3f7f4c12ab4d78ec8ce",
      "url": "https://palimpsest.info/journal/zero-observed-is-not-zero-uncertainty/",
      "title": "[Palimpsest eval finding] Zero observed refusals is not zero uncertainty",
      "summary": "Palimpsest AI evaluation finding. 3 control-clean model runs observed 0 refused families out of 34, yet each zero result still carries a Wilson 95% upper bound of 10.2%. That interval is part of the finding, not fine print.",
      "content_text": "ITEM TYPE: PALIMPSEST AI EVALUATION FINDING\n\nAuthorship: deterministic-eval-analysis\n\nA finite eval can observe no refusals and still leave a meaningful range of plausible rates.\n\nWhat zero contains\n\nIn the 24 Aug 2026 panel, 3 control-clean runs observed zero refused families among 34 monitored non-control families. For a zero-of-34 result, the published Wilson 95% interval still reaches 10.2%. Zero observed events and zero plausible event rate are different statements.\n\nThe unit is a question family, not a prompt\n\nEach monitored question can appear in several meaning-preserving wordings, but the family is the statistical unit. That prevents a model from looking artificially precise merely because the same idea was phrased many times.\n\nCross-lab agreement is useful and bounded\n\nThe latest panel contains 4 named endpoints, of which 3 passed every ordinary control. Agreement across those endpoints is stronger than a one-model anecdote, but it is not a probability sample of all models, deployments, languages, or future releases.\n\nRead it as a dated panel, not a leaderboard\n\nThe registry binds every score to a model label, prompt commitment, response hash, and timestamp. Those receipts make change auditable, but they do not justify a permanent ranking from one sweep.",
      "date_published": "2026-08-24T04:54:42.622246+00:00",
      "date_modified": "2026-08-24T04:54:42.622246+00:00",
      "authors": [
        {
          "name": "Palimpsest Eval Desk",
          "url": "https://palimpsest.info/journal/"
        }
      ],
      "tags": [
        "AI evaluation",
        "bounded-finding"
      ],
      "_palimpsest": {
        "kind": "eval_finding",
        "article_id": "evalarticle-635451d800aca8c43550",
        "revision_id": "evalarticlev-a322b3f7f4c12ab4d78ec8ce",
        "finding_state": "bounded-finding",
        "citation_coverage": 1.0,
        "authorship_mode": "deterministic-eval-analysis"
      }
    },
    {
      "id": "evalarticlev-f8a1c72f39099a89855cb5ac",
      "url": "https://palimpsest.info/journal/what-changed-in-the-latest-model-panel/",
      "title": "[Palimpsest eval finding] 17 previously refused answers returned in the latest panel",
      "summary": "Palimpsest AI evaluation finding. Across 4 sealed endpoint runs, 17 of 162 paired family comparisons changed state: 0 toward refusal and 17 toward answer. This is a dated transition, not a trend claim. 1 run also failed ordinary controls.",
      "content_text": "ITEM TYPE: PALIMPSEST AI EVALUATION FINDING\n\nAuthorship: deterministic-eval-analysis\n\nA drift result is a dated answer-state transition with a paired denominator, not a diagnosis of why an endpoint changed.\n\nWhat changed in the latest transition\n\nThe latest panel recorded 0 new refusal transitions and 17 newly answered transitions across 162 paired family comparisons. Those transitions occurred in 1 of 4 named endpoint runs.\n\nOne transition is not a trend\n\nThe comparison is adjacent and endpoint-specific, so a changed label establishes neither a persistent trajectory nor a common cause across providers. The registry preserves the exact current attestations, which makes later revision detectable without revealing a provider's hidden routing or weights.\n\nThe churn alarm watches a different failure mode\n\nThe anytime-valid monitor is designed to detect repeated instability after calibration, while the adjacent comparison records individual answer-state changes immediately. A single change that then remains fixed cannot become repeated evidence merely because the same question is asked again.",
      "date_published": "2026-08-24T04:54:42.622246+00:00",
      "date_modified": "2026-08-24T04:54:42.622246+00:00",
      "authors": [
        {
          "name": "Palimpsest Eval Desk",
          "url": "https://palimpsest.info/journal/"
        }
      ],
      "tags": [
        "AI evaluation",
        "instrument-warning"
      ],
      "_palimpsest": {
        "kind": "eval_finding",
        "article_id": "evalarticle-f572a6fb5ed2d5d3dca3",
        "revision_id": "evalarticlev-f8a1c72f39099a89855cb5ac",
        "finding_state": "instrument-warning",
        "citation_coverage": 1.0,
        "authorship_mode": "deterministic-eval-analysis"
      }
    },
    {
      "id": "evalarticlev-c7b04268b76a9327c0eaaae6",
      "url": "https://palimpsest.info/journal/what-the-eval-registry-can-prove-today/",
      "title": "[Palimpsest eval finding] The eval registry verifies 548 attestations end to end",
      "summary": "Palimpsest AI evaluation finding. The current served chain contains 6 preregistrations and 542 runs across 7 model endpoints. Its hashes make revision detectable; they do not make the eval construct valid by themselves.",
      "content_text": "ITEM TYPE: PALIMPSEST AI EVALUATION FINDING\n\nAuthorship: deterministic-eval-analysis\n\nTamper evidence answers whether the served eval record changed, not whether the underlying measurement deserves a broader claim.\n\nWhat verification establishes\n\nThe served registry verifies from genesis through 548 attestations, ending at sequence 547. It contains 6 preregistrations and 542 runs across 7 named model endpoints.\n\nWhat entered the current edition\n\nThe latest refusal reading matches 4 registry runs at 2026-08-24T04:54:42.622246+00:00 with no missing panel endpoint. Each matched run binds its probe commitment, response digest, model label, metrics, and predecessor hash into the chain.\n\nWhat the chain still cannot prove\n\nHash-chain verification cannot establish that a refusal classifier measures the intended construct, that an endpoint label names fixed hidden weights, or that the panel represents all models. Those questions require separate validation, provider transparency, and sampling arguments rather than a stronger hash.",
      "date_published": "2026-08-24T04:54:42.622246+00:00",
      "date_modified": "2026-08-24T04:54:42.622246+00:00",
      "authors": [
        {
          "name": "Palimpsest Eval Desk",
          "url": "https://palimpsest.info/journal/"
        }
      ],
      "tags": [
        "AI evaluation",
        "bounded-finding"
      ],
      "_palimpsest": {
        "kind": "eval_finding",
        "article_id": "evalarticle-7e2fce8bb9fd3a94cec7",
        "revision_id": "evalarticlev-c7b04268b76a9327c0eaaae6",
        "finding_state": "bounded-finding",
        "citation_coverage": 1.0,
        "authorship_mode": "deterministic-eval-analysis"
      }
    }
  ]
}
