How well can we tell a tampered model from a clean one?
How well can we tell a tampered model from a clean one? We run the read on the public benchmarks security researchers use — and post every score, with its uncertainty, and the misses right beside the catches.
Graded on the same public exams researchers use — NIST TrojAI · PADBench · BAIT-ModelZoo. See the specimens behind every score →
What about a pool that’s mostly clean?
The fair test of any detector: run it on a pool that is almost all clean — the real-world ratio — and count the false alarms. We publish that number here, on a channel that needs no matched reference, including the half that doesn’t flatter us.
Zero false alarms on the non-collision pool — 192 benign adapters, every one scored on this channel except the deliberately hardest. An ordinary benign adapter has little to wake. It is not perfectly silent, though: on the hardest negatives — benign finetunes whose legitimate job is the payload’s own shape — 1 of 18 confessed. Pooled with those, each adapter counted exactly once, the honest number is 1 in 207, disclosed. The ceiling is the Wilson 95% bound — the same figure the runnable check below prints, so the page, the tool, and the reply all agree (the exact one-sided Clopper–Pearson is tighter still).
We withdrew two denominators — and they were withdrawn in our favour.
No measurement moved. Every fire count on this panel is the number the artifacts always held. What was wrong was two TOTALS that could not be re-derived from their parts, and a total you cannot re-derive is not evidence. Both are withdrawn rather than restated — and one sentence that was simply false is struck below by name rather than quietly rewritten. The importer now refuses any figure this record retracts.
- 0/68 Real cohorts (0/48 same-recipe + 0/20 replayable cross-recipe), but published as a TOTAL while the 144-adapter community pool — measured 2026-07-23, the day before publication — sat outside it. It names 68 of 192. Withdrawn, not repaired.
- 1/210 48 + 144 + 18 summed WITHOUT subtracting the 3 adapters that belong to both the community pool and the hardest-negative cohort, so it counted no set of adapters that exists. The correct pooled denominator is 207. Withdrawn, not repaired.
- An earlier revision of this file asserted the fire was 'the single own-task-collision fire over the whole 210-adapter community pool the weight-read fired 210/210 on.' That sentence was FALSE: gcp-acp/retailproducts-llama-gemma-2-9b-it-lora is not a member of that 210-adapter cohort. Retracted here rather than silently rewritten.
- 210 ALSO names a real and correct figure elsewhere: the community cohort the ABSOLUTE weight-novelty read fires 210/210 on (features_index.json, 210 entries; RESULTS.json community FPR 1.0). That 210/210 stands. It is a DIFFERENT SET from any elicitation denominator — 15 of the 18 payload-shaped hardest negatives, the one that fired included, are absent from it. Two sets happening to number 210 is exactly why the pooled 210 is withdrawn rather than repaired.
The correction runs in our favour: we had understated our own clean evidence by nearly threefold. That is exactly why it needed doing carefully, and why the parts are published rather than a new total.
How every bound above is recomputed
Every bound above is recomputable from (flags, n) alone. Clopper-Pearson upper U solves sum_{k=0..x} C(n,k) U^k (1-U)^(n-k) = alpha, with alpha = 0.05/0.01 one-sided and 0.025/0.005 two-sided; for x=0 that closes to U = 1 - alpha^(1/n). Wilson two-sided 95% upper uses z=1.96: ((p + z^2/2n) + z*sqrt(p(1-p)/n + z^2/4n^2)) / (1 + z^2/n). Rule of three = 3/n. Values rounded to 5 decimals. Rerunning that code on the three cohorts published 2026-07-24 (0/68, 1/18, 1/210) reproduces their stated bounds exactly, which is how this recomputation was validated before the denominators moved.
Priced on 1/207 — a 2.7% false-alarm ceiling, wider than the 2.0% the corrected non-collision denominator earns. Priced on the pooled cohort — the one that carries our own false fire — not on the non-collision pool. The correction made the non-collision denominator larger and its ceiling tighter, and pricing this row there would have improved every number on it without a single new measurement.
A near-zero point estimate is not a guarantee, and at a real-world base rate a single fire is only right in the low single digits. So this is not a solved detector — it is a channel whose confession we trust and whose silence we don’t.
And how often does it fire when it should? Everything above is the false-alarm side. This is the other half, and it is the weaker one: 38% pooled recall on one poison subclass and 46% on the other — between 12% and 75% depending on the architecture.
We publish both numbers rather than an average, because the average is not a measurement of anything. This channel is precision-optimised: when it fires you can act on it, and it wakes only a minority of poisoned models. That is why silence here is recorded as not attested and never as clean — a majority of real tampering does not wake it, so reading silence as a clearance would be wrong most of the time it mattered. Poison recall is capped at n=8 per subclass per architecture, so these confidence intervals are wide.
The 2 assumptions this ceiling rests on — and where it doesn’t transfer
- Exchangeability. The bound is distribution-free but NOT assumption-free. It holds under the single assumption that the clean calibration cohort is EXCHANGEABLE with the clean inputs seen at deployment (drawn from the same population of benign finetunes). If real-world clean adapters differ systematically from these 192 (novel recipes, architectures, or task families outside the tested set), the guarantee does not transfer and the FPR must be re-estimated on an in-distribution clean pool. The payload-shaped hardest-negative cohort is the deliberate stress of exactly that assumption, and it is where the one fire came from — the honest reading of how far the non-collision bound transfers to a collision-rich deployment.
- Independence. The Clopper-Pearson bound is a binomial statement assuming INDEPENDENT Bernoulli draws. Several cohort members are author/base-clustered: the replayable community subset of 20 is only 14 distinct authors (Jazhyc x3, Jongbin-kr x3, ArchSid x2, HermitQ x2 sibling variants), and the 48 same-recipe controls are recipe families replicated across 3 architectures. Positive within-cluster correlation shrinks effective n below nominal, so the true upper bounds are modestly WIDER than the nominal values quoted here. No cluster-corrected interval is computed: the non-collision data supplies no intra-cluster correlation estimate at zero fires, and the full 144-adapter pool's clustering has not been measured. The direction and the effective-n counts we do have are disclosed instead.
This FPR is the presence-by-confession (reference-model-free payload elicitation) channel, offered IN PLACE OF the weight-read presence axis (which remains FPR=1.0, conceded/published). It is a separate signal, not a calibration of the failed weight-read presence axis.
The one door that answers a stranger in JSON.
One URL, one question — Does Vulcora hold a signed attestation dossier for this model? The answer is either the signed dossier itself or one of 8 named reasons we serve none. The vocabulary is closed on purpose: a refusal that conflates “we never looked” with “we looked and are withholding the verdict” — or with “we hold signed evidence that contradicts what we publish” — cannot be audited, so ours doesn’t.
https://api.vulcora.se/api/attestationsthe whole plane — every model it will serve a dossier for, computed by the same function the per-model route answers with, so the list and the answers cannot drift apartGEThttps://api.vulcora.se/api/attestations/{owner}/{model}one model — worked example: Qwen/Qwen2.5-0.5B-Instructcomm -13 <(curl -s https://api.vulcora.se/api/attestations | jq -r '.attested[]' | sort) \
<(curl -s https://api.vulcora.se/api/model_records | jq -r '.data[] | select(.attributes.protora_assessed) | .id' | sort)Every line it prints is a model we publish a verdict about and hold no signed dossier behind. We are not asking you to believe the gap is small; we are handing you the command that measures it. Nothing in it trusts this page — both halves come straight off the plane, and the subtraction happens on your side.
We could not reach the plane just now, so no coverage number is printed here — a figure from an earlier read, shown as if it were current, is precisely the kind of claim this page exists to refuse. Run the command; its answer is the one that counts.
The public keys are served by the same host that serves the dossiers, so the trust anchor is circular. A passing check proves the dossier is byte-for-byte the one signed by the holder of the named key, and that its verdict was not altered in transit or at rest — not that the key is ours. It cannot rule out a fully compromised host serving a self-consistent forgery. Until that key is mirrored in a channel independent of this API, treat a passing check as unanchored. Pin the key id and bytes now, obtain them from a second channel later, and re-check against your pinned copy.
The plane volunteers this itself, unprompted, in verify.key_provenance. It is restated here because a page that is less honest than its own API is a regression.
The 8 named reasons the plane serves no dossier — and what each one concedes. Not all of them mean we hold nothing.
unknown_modelverdict_under_reviewnot_assessedno_signed_dossierattestation_subject_mismatchattestation_verdict_mismatchattestation_tier_conflictattestation_verdict_unsigned
And one outcome that is not a negative at all: dossier_not_firewall_clean — a 500, meaning a dossier we hold failed our own read-time content gate. That is a bug on our side, and we refuse to serve the record rather than serve it unchecked. It is not a statement that we hold nothing.
The bodies are served by the plane and are not restated from memory here — open the index to read them.
Each score is the chance the read ranks a tampered model as more suspicious than a clean one. 1.00 is perfect. 0.50 is a coin-flip — no signal at all.
Every score we can defend — with its uncertainty.
The read has three regimes, and the honest picture needs all three. One row per benchmark below — the bar is the score, with its uncertainty and sample size beside it.
The misses, on the same page as the wins.
A read that only ever wins is lying. These inverted, went not-applicable, or returned an honest null — so we publish them as exactly that.
Surveyed, no honest offline path: NIST TrojAI llm-pretrain-apr2024 (Llama-2-7B, call-and-response) · BackdoorLLM (bboylyg) — 13B/70B + 5 DPA families · TDC 2023 (NeurIPS Trojan Detection Challenge).
The limits ride with every score.
Stated in full, because the honesty is the credibility. Straight from the record.
Read all 11 limits — straight from the record
- "reference-MODEL-free", never bare "reference-free" — state the matched-reference requirement EVERY time.
- Absolute-deploy concession: no matched ref → finetuning-distribution detector, ~100% FPR.
- NOT first / NOT best. PEFTGuard reports ~1.0 per-family on PADBench with a trained meta-classifier; our operating point is ref-MODEL-free + forward-free + CPU + no per-family refit. Coverage + honesty point, not a SOTA/ranking claim.
- Distant-arch / low-rank drop is REAL: RoBERTa 0.791, r16 0.719 vs VLM/LoRA+/cba 0.94–1.0. State the ceiling with every headline.
- Small-n CIs. BAIT llama2 n=13: AUC 1.0 but boot CI [1.0,1.0] is a DEGENERATE artifact of perfect separation, not a certainty claim — the rigorous null is the exhaustive label-permutation perm_p=0.0012 (C(13,6)=1716). Per-cohort 30/30 → wide CIs (some span 0.57–0.96).
- Single attack family per corpus. BAIT = cba only → measures cba separability, not "any backdoor" / "always detectable."
- Full-FT substrate NOT yet earned — currently inverting at n=6; do not report it as coverage.
- Public artifacts = AUC / FPR / recall@FPR / CI / n ONLY. No sealed-method tokens.
- Specificity discipline works BOTH ways. Travis-ML n_clean=1 + seed-confounded design → the read returns an honest NULL (ρ=0.038), NOT the degenerate AUC=1.0 the data would allow. Do not force a single-clean/seed-confounded design to a number.
- Balance the reference before quoting a paired-LOO number. thoughtworks std-LOO AUC=1.0 is a reference-SIZE artifact (poison scored vs 8 clean, clean scored vs 7). The balanced read (0.781) removes the asymmetry — always balance before quoting.
- BAIT-ModelZoo epoch-confound — CONCEDE standalone. Every clean adapter was trained 2 epochs, every poison 4, so "backdoor present" and "trained 2× longer" are one variable (epoch⊥label collinearity r = 1.000, zero duration-matched clean controls). The read is magnitude-invariant so normalization is a no-op (ΔAUC ~1.5e-7); regressing the epoch covariate out collapses Mistral to chance (0.50) — you cannot partial out a column identical to the answer. So `perm_p` certifies only that SOME structure separates the two groups, NOT that the structure is the backdoor. Post the AUC 1.0 ONLY with this guard on the same line; lead the trained class with the parity-clean NIST TrojAI-96 (0.878, every training knob matched), never with BAIT-ModelZoo.
Not a scoreboard — a proof.
Each read ships with a per-model dossier: the real per-scanner board, and a machine-checked proof you can re-verify yourself.
In technical terms: reference-MODEL-free, forward-free, CPU-only weight-space backdoor-PRESENCE read over LoRA / full-FT weight deltas.
Source — the engine coverage matrix, 2026-07-23. Public numbers are AUC / FPR / CI / n only.