We hid a backdoor in a public model — and sealed the answer first

In one breathWe published the answer's fingerprint before opening the challenge — so nobody, us included, can quietly move the goalposts.
For a stretch this month there were seven open AI models sitting in public, all published by us — and one of them was lying. Anyone could download them with ordinary tooling and try to answer one question: which one betrays you, and what's the secret word that wakes it?
The part we care about most isn't the trap. It's that we sealed the answer before opening the doors — we published a short fingerprint of the solution first, so we couldn't quietly change our story later. That's the whole discipline of this house in one move: <span class="vb-u">prove it in a way a stranger can check, or don't say it.</span>
<p class="vb-inline-cta">The full, replayable record lives here → <a href="/record">The record</a></p>Provenance — verified in your browser
The bytes of this post were hashed in your browser and match the value bound to a public, dated commit. You did not have to trust us — you just re-checked.
- content hash
f7ad08bcf122…d388bc53sha256 of these bytes = signed content_sha256- git blob
b2f9391a8de1…4b317857equals `git hash-object` on the public file- signature
Ed25519 · key 1f9d…7cbVulcora signed this ledger entry- record
entry 20 · cf532e91fc…bd4a31published 2026-07-27
Check it yourself, no code of ours required: git hash-object posts/home/the-sealed-challenge.md in a clone of the public post repository must print the git blob above. This binding is tamper-evident, not tamper-proof — GitHub is a trusted third party, not a mathematical guarantee, and we would rather say so than imply something stronger.