<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Palimpsest AI eval results</title><link>https://palimpsest.info/journal/</link><description>Palimpsest's own dated AI evaluation results, with controls, uncertainty, deterministic authorship and revision identity attached.</description><language>en</language><lastBuildDate>Mon, 24 Aug 2026 04:54:42 +0000</lastBuildDate><atom:link href="https://palimpsest.info/journal/feed.xml" rel="self" type="application/rss+xml" /><item><title>[Palimpsest eval finding] The headline was 0%. The controls still failed.</title><link>https://palimpsest.info/journal/before-reading-the-score-read-the-controls/</link><guid isPermaLink="false">evalarticlev-7dff12f8c34e0a64e9910063</guid><pubDate>Mon, 24 Aug 2026 04:54:42 +0000</pubDate><description>Palimpsest AI evaluation finding. Meta Llama 3.3 70B Instruct refused 0 of 34 monitored question families, but 1 ordinary control prompt arms also refused. That makes the result an instrument warning, not a censorship finding. Authorship: deterministic-eval-analysis.</description><category>palimpsest-eval-finding</category></item><item><title>[Palimpsest eval finding] Zero observed refusals is not zero uncertainty</title><link>https://palimpsest.info/journal/zero-observed-is-not-zero-uncertainty/</link><guid isPermaLink="false">evalarticlev-a322b3f7f4c12ab4d78ec8ce</guid><pubDate>Mon, 24 Aug 2026 04:54:42 +0000</pubDate><description>Palimpsest AI evaluation finding. 3 control-clean model runs observed 0 refused families out of 34, yet each zero result still carries a Wilson 95% upper bound of 10.2%. That interval is part of the finding, not fine print. Authorship: deterministic-eval-analysis.</description><category>palimpsest-eval-finding</category></item><item><title>[Palimpsest eval finding] 17 previously refused answers returned in the latest panel</title><link>https://palimpsest.info/journal/what-changed-in-the-latest-model-panel/</link><guid isPermaLink="false">evalarticlev-f8a1c72f39099a89855cb5ac</guid><pubDate>Mon, 24 Aug 2026 04:54:42 +0000</pubDate><description>Palimpsest AI evaluation finding. Across 4 sealed endpoint runs, 17 of 162 paired family comparisons changed state: 0 toward refusal and 17 toward answer. This is a dated transition, not a trend claim. 1 run also failed ordinary controls. Authorship: deterministic-eval-analysis.</description><category>palimpsest-eval-finding</category></item><item><title>[Palimpsest eval finding] The eval registry verifies 548 attestations end to end</title><link>https://palimpsest.info/journal/what-the-eval-registry-can-prove-today/</link><guid isPermaLink="false">evalarticlev-c7b04268b76a9327c0eaaae6</guid><pubDate>Mon, 24 Aug 2026 04:54:42 +0000</pubDate><description>Palimpsest AI evaluation finding. The current served chain contains 6 preregistrations and 542 runs across 7 model endpoints. Its hashes make revision detectable; they do not make the eval construct valid by themselves. Authorship: deterministic-eval-analysis.</description><category>palimpsest-eval-finding</category></item></channel></rss>
