How well can we tell a tampered model from a clean one?
How well can we tell a tampered model from a clean one? We run the read on the public benchmarks security researchers use — and post every score, with its uncertainty, and the misses right beside the catches.
Graded on the same public exams researchers use — PADBench · NIST TrojAI · BAIT-ModelZoo. The plain-language head-to-head →
Each score is the chance the read ranks a tampered model as more suspicious than a clean one. 1.00 is perfect. 0.50 is a coin-flip — no signal at all.
Every score we can defend — with its uncertainty.
The read has three regimes, and the honest picture needs all three. One row per benchmark below — the bar is the score, with its uncertainty and sample size beside it.
The misses, on the same page as the wins.
A read that only ever wins is lying. These inverted, went not-applicable, or returned an honest null — so we publish them as exactly that.
Surveyed, no honest offline path: NIST TrojAI llm-pretrain-apr2024 (Llama-2-7B, call-and-response) · BackdoorLLM (bboylyg) — 13B/70B + 5 DPA families · TDC 2023 (NeurIPS Trojan Detection Challenge).
The limits ride with every score.
Stated in full, because the honesty is the credibility. Straight from the record.
Read all 11 limits — straight from the record
- "reference-MODEL-free", never bare "reference-free" — state the matched-reference requirement EVERY time.
- Absolute-deploy concession: no matched ref → finetuning-distribution detector, ~100% FPR.
- NOT first / NOT best. PEFTGuard reports ~1.0 per-family on PADBench with a trained meta-classifier; our operating point is ref-MODEL-free + forward-free + CPU + no per-family refit. Coverage + honesty point, not a SOTA/ranking claim.
- Distant-arch / low-rank drop is REAL: RoBERTa 0.791, r16 0.719 vs VLM/LoRA+/cba 0.94–1.0. State the ceiling with every headline.
- Small-n CIs. BAIT llama2 n=13: AUC 1.0 but boot CI [1.0,1.0] is a DEGENERATE artifact of perfect separation, not a certainty claim — the rigorous null is the exhaustive label-permutation perm_p=0.0012 (C(13,6)=1716). Per-cohort 30/30 → wide CIs (some span 0.57–0.96).
- Single attack family per corpus. BAIT = cba only → measures cba separability, not "any backdoor" / "always detectable."
- Full-FT substrate NOT yet earned — currently inverting at n=6; do not report it as coverage.
- Public artifacts = AUC / FPR / recall@FPR / CI / n ONLY. No sealed-method tokens.
- Specificity discipline works BOTH ways. Travis-ML n_clean=1 + seed-confounded design → the read returns an honest NULL (ρ=0.038), NOT the degenerate AUC=1.0 the data would allow. Do not force a single-clean/seed-confounded design to a number.
- Balance the reference before quoting a paired-LOO number. thoughtworks std-LOO AUC=1.0 is a reference-SIZE artifact (poison scored vs 8 clean, clean scored vs 7). The balanced read (0.781) removes the asymmetry — always balance before quoting.
- BAIT-ModelZoo epoch-confound — CONCEDE standalone. Every clean adapter was trained 2 epochs, every poison 4, so "backdoor present" and "trained 2× longer" are one variable (epoch⊥label collinearity r = 1.000, zero duration-matched clean controls). The read is magnitude-invariant so normalization is a no-op (ΔAUC ~1.5e-7); regressing the epoch covariate out collapses Mistral to chance (0.50) — you cannot partial out a column identical to the answer. So `perm_p` certifies only that SOME structure separates the two groups, NOT that the structure is the backdoor. Post the AUC 1.0 ONLY with this guard on the same line; lead the trained class with the parity-clean NIST TrojAI-96 (0.878, every training knob matched), never with BAIT-ModelZoo.
What about a pool that’s mostly clean?
The fair test of any detector: run it on a pool that is almost all clean — the real-world ratio — and count the false alarms. We publish that number here, on a channel that needs no matched reference, including the half that doesn’t flatter us.
Zero false alarms on the non-collision community pool — an ordinary benign adapter has little to wake. It is not perfectly silent, though: on the hardest negatives — benign finetunes whose legitimate job is the payload’s own shape — one of eighteen confessed, so the honest pooled number is 1 in 210, disclosed. The ceiling is the Wilson 95% bound — the same figure the runnable check below prints, so the page, the tool, and the reply all agree (the exact one-sided Clopper–Pearson is tighter still).
A near-zero point estimate is not a guarantee, and at a real-world base rate a single fire is only right in the low single digits. So this is not a solved detector — it is a channel whose confession we trust and whose silence we don’t.
The 2 assumptions this ceiling rests on — and where it doesn’t transfer
- Exchangeability. The bound is distribution-free but NOT assumption-free. It holds under the single assumption that the clean calibration cohort is EXCHANGEABLE with the clean inputs seen at deployment (drawn from the same population of benign finetunes). If real-world clean adapters differ systematically from these 68 (novel recipes, architectures, or task families outside the tested set), the guarantee does not transfer and the FPR must be re-estimated on an in-distribution clean pool. n is modest (20 cross-recipe), so the cross-recipe upper bound is the widest and the load-bearing one.
- Independence. The Clopper-Pearson bound is a binomial statement assuming INDEPENDENT Bernoulli draws. Several cohort members are author/base-clustered: cross-recipe n=20 is only 14 distinct authors (Jazhyc x3, Jongbin-kr x3, ArchSid x2, HermitQ x2 sibling variants), and same-recipe n=48 is 16 recipes replicated across 3 architectures. Positive within-cluster correlation shrinks effective n below nominal, so the true upper bounds are modestly WIDER than the nominal values quoted here. No cluster-corrected interval is computed because the all-zero data supplies no intra-cluster correlation estimate; the direction and the effective-n counts are disclosed instead.
This FPR is the presence-by-confession (reference-model-free payload elicitation) channel, offered IN PLACE OF the weight-read presence axis (which remains FPR=1.0, conceded/published). It is a separate signal, not a calibration of the failed weight-read presence axis.
Not a scoreboard — a proof.
Each read ships with a per-model dossier: the real per-scanner board, and a machine-checked proof you can re-verify yourself.
In technical terms: reference-MODEL-free, forward-free, CPU-only weight-space backdoor-PRESENCE read over LoRA / full-FT weight deltas.
Source — the engine coverage matrix, 2026-07-23. Public numbers are AUC / FPR / CI / n only.