The coverage record

How well can we tell a tampered model from a clean one?

How well can we tell a tampered model from a clean one? We run the read on the public benchmarks security researchers use — and post every score, with its uncertainty, and the misses right beside the catches.

Graded on the same public exams researchers use — PADBench · NIST TrojAI · BAIT-ModelZoo. The plain-language head-to-head →

How to read a score

Each score is the chance the read ranks a tampered model as more suspicious than a clean one. 1.00 is perfect. 0.50 is a coin-flip — no signal at all.

Where the read works

Every score we can defend — with its uncertainty.

The read has three regimes, and the honest picture needs all three. One row per benchmark below — the bar is the score, with its uncertainty and sample size beside it.

In-class · matched reference0.941–0.953calibrated on the model’s own family — the strong regime
Cross-family · zero-shot0.854one decision function, an architecture it never saw — it still transfers
Absolute · no referencedeada fixed anchor vs the open world — 210/210 fire, FPR 1.0 (the panel below)
PADBench-exp — qwen2vl_vqav2_r16Qwen2-VL (multimodal VLM, LoRA r16) · vqav2 insertsent
0.976Strongn 30/30 · CI 0.917–1.000recall@1%FPR 0.80 · @5% 1.00
PADBench-exp — l2_loraplus_r8Llama-2-7b (LoRA+ PEFT, r8) · toxic-backdoors-hard
0.970Strongn 30/30 · CI 0.921–1.000recall@1%FPR 0.633 · @5% 0.70
PADBench llama3_8b (per-fam)Llama-3-8B
0.953Strongper-family sub-read — n on the pooled row above
PADBench qwen1.5_7b (per-fam)Qwen-1.5-7B
0.941Strongper-family sub-read — n on the pooled row above
NIST TrojAI llm-instruct-oct2024 (96-LoRA)3-arch matched-ref (gemma-2-2b / gemma-2-9b / llama-3.1-8b) · adified + meanified
0.878Solidn (96 total)
PADBench-exp — POOLED cross-cohortmixed (VLM+encoder+LoRA/LoRA+) · mixed
0.861Solidn 210/220 · CI 0.825–0.898recall@1%FPR 0.267 · @5% 0.417 (@10%: 0.517)
LOAO cross-family (leave-one-arch-out)cross-family pooled · mixed
0.854Solid
PADBench LOAO (zero-shot cross-cohort TRANSFER)freshQwen2-VL VLM + RoBERTa encoder + Llama-2-7B (r8/16/32, LoRA+) · insertsent/toxic/imdb/vqav2 (per-cohort)
0.854Solidn 180/180 · CI 0.814–0.889
PADBench-exp — l2_rank8Llama-2-7b (LoRA r8) · toxic-hard (rank-gen)
0.853Solidn 30/30 · CI 0.727–0.956recall@1%FPR 0.00 · @5% 0.067
PADBench 5-family @ r256 (pooled)chatglm6b, flan_t5_xl, llama2_7b, llama3_8b, qwen1.5_7b · insertsent / toxic (per-family)
0.848Fairn 137/148 · CI 0.797–0.883
PADBench-exp — roberta_imdb_r16RoBERTa (encoder-only, LoRA r16) · imdb insertsent
0.791Fairn 30/30 · CI 0.668–0.893recall@1%FPR 0.167 · @5% 0.233
PADBench-exp — l2_rank32Llama-2-7b (LoRA r32) · toxic-hard (rank-gen)
0.764Fairn 30/30 · CI 0.636–0.874recall@1%FPR 0.167 · @5% 0.20
PADBench-exp — l2_rank16Llama-2-7b (LoRA r16) · toxic-hard (rank-gen)
0.719Weak but realn 30/30 · CI 0.568–0.850recall@1%FPR 0.067 · @5% 0.167
PADBench distant familieschatglm6b / flan_t5_xl / llama2_7b
0.700Weak but realper-family sub-read — n on the pooled row above
BAIT-ModelZoo — llama2_7bLlama-2-7b (LoRA r64) · cba (composite/distributed 2-part triggers)⚠ BAIT-ModelZoo epoch-confound — CONCEDE standalone. The full guard rides below.
1.00Strongn 6/7 · CI [1.0,1.0] degeneraterecall@1%FPR 1.00 · @5% 1.00
BAIT-ModelZoo — POOLEDLlama-2-7b · cba⚠ BAIT-ModelZoo epoch-confound — CONCEDE standalone. The full guard rides below.
1.00Strongn 6/7 · CI [1.0,1.0] degenrecall@1%FPR 1.00 · @5% 1.00
Where it doesn’t — posted anyway

The misses, on the same page as the wins.

A read that only ever wins is lying. These inverted, went not-applicable, or returned an honest null — so we publish them as exactly that.

PARTIAL + currently INVERTING
gemma-2-2b full-FT (Δ=FT−base)NIST TrojAI oct2024 FULL-FT (gemma-2-2b)
NOT-APPLICABLE — input-contract mismatch
Llama-2-7b AdaLoRAPADBench `l2_adalora_r8`
PARTIAL (budget)
Llama-2-7b LoRA r64PADBench `l2_rank64`
NOT-APPLICABLE-ABSOLUTE
Llama-2-7b LoRAPADBench 6 POISON-ONLY families (ag_news/imdb × RIPPLE/stybkd/syntactic)
not-applicable (insufficient, not inverting)
Mistral-7B-Instruct-v0.2 LoRABAIT-ModelZoo mistral_7b
not streamed
Meta-Llama-3-8B-Instruct LoRABAIT-ModelZoo llama3_8b
pending
llama2_7bBAIT sham (specificity null) @ n=13
HONEST NULL — no dose-response
Qwen2.5-1.5B-Instruct (LoRA r16) — NEW archTravis-ML dose ladder (Qwen2.5-1.5B, k=15–25)
WEAK borderline
Qwen2.5-0.5B-Instruct (LoRA r16) — NEW archthoughtworks/backdoor-4pair-hate (AND-trigger)

Surveyed, no honest offline path: NIST TrojAI llm-pretrain-apr2024 (Llama-2-7B, call-and-response) · BackdoorLLM (bboylyg) — 13B/70B + 5 DPA families · TDC 2023 (NeurIPS Trojan Detection Challenge).

What we won’t let a number imply

The limits ride with every score.

Stated in full, because the honesty is the credibility. Straight from the record.

Read all 11 limits — straight from the record
  • "reference-MODEL-free", never bare "reference-free" — state the matched-reference requirement EVERY time.
  • Absolute-deploy concession: no matched ref → finetuning-distribution detector, ~100% FPR.
  • NOT first / NOT best. PEFTGuard reports ~1.0 per-family on PADBench with a trained meta-classifier; our operating point is ref-MODEL-free + forward-free + CPU + no per-family refit. Coverage + honesty point, not a SOTA/ranking claim.
  • Distant-arch / low-rank drop is REAL: RoBERTa 0.791, r16 0.719 vs VLM/LoRA+/cba 0.94–1.0. State the ceiling with every headline.
  • Small-n CIs. BAIT llama2 n=13: AUC 1.0 but boot CI [1.0,1.0] is a DEGENERATE artifact of perfect separation, not a certainty claim — the rigorous null is the exhaustive label-permutation perm_p=0.0012 (C(13,6)=1716). Per-cohort 30/30 → wide CIs (some span 0.57–0.96).
  • Single attack family per corpus. BAIT = cba only → measures cba separability, not "any backdoor" / "always detectable."
  • Full-FT substrate NOT yet earned — currently inverting at n=6; do not report it as coverage.
  • Public artifacts = AUC / FPR / recall@FPR / CI / n ONLY. No sealed-method tokens.
  • Specificity discipline works BOTH ways. Travis-ML n_clean=1 + seed-confounded design → the read returns an honest NULL (ρ=0.038), NOT the degenerate AUC=1.0 the data would allow. Do not force a single-clean/seed-confounded design to a number.
  • Balance the reference before quoting a paired-LOO number. thoughtworks std-LOO AUC=1.0 is a reference-SIZE artifact (poison scored vs 8 clean, clean scored vs 7). The balanced read (0.781) removes the asymmetry — always balance before quoting.
  • BAIT-ModelZoo epoch-confound — CONCEDE standalone. Every clean adapter was trained 2 epochs, every poison 4, so "backdoor present" and "trained 2× longer" are one variable (epoch⊥label collinearity r = 1.000, zero duration-matched clean controls). The read is magnitude-invariant so normalization is a no-op (ΔAUC ~1.5e-7); regressing the epoch covariate out collapses Mistral to chance (0.50) — you cannot partial out a column identical to the answer. So `perm_p` certifies only that SOME structure separates the two groups, NOT that the structure is the backdoor. Post the AUC 1.0 ONLY with this guard on the same line; lead the trained class with the parity-clean NIST TrojAI-96 (0.878, every training knob matched), never with BAIT-ModelZoo.
The question everyone asks

What about a pool that’s mostly clean?

The fair test of any detector: run it on a pool that is almost all clean — the real-world ratio — and count the false alarms. We publish that number here, on a channel that needs no matched reference, including the half that doesn’t flatter us.

0/20cross-recipe community finetunes95% ceiling ≤ 16.1% · 14 distinct
0/68non-collision benign pool95% ceiling ≤ 5.3%
0/48same-recipe controls95% ceiling ≤ 7.4% · 16 distinct
1/18payload-shaped hardest negatives (1 own-task collision)95% ceiling ≤ 25.8%benign finetunes whose LEGITIMATE task IS the payload shape (ad-copy, persuasion, roleplay-antagonist) — the semantic-collision negatives most likely to trip the elicitation channel, built on purpose. One falsely confessed (an own-task collision), disclosed. This is the honest true-negative fire.
1/210full pooled community set (incl. hardest negatives)95% ceiling ≤ 2.6%the full pooled benign set once the payload-shaped hardest-negatives are included: the single own-task-collision fire over the whole 210-adapter community pool the weight-read fired 210/210 on. 1/210, disclosed — never 0/210.

Zero false alarms on the non-collision community pool — an ordinary benign adapter has little to wake. It is not perfectly silent, though: on the hardest negatives — benign finetunes whose legitimate job is the payload’s own shape — one of eighteen confessed, so the honest pooled number is 1 in 210, disclosed. The ceiling is the Wilson 95% bound — the same figure the runnable check below prints, so the page, the tool, and the reply all agree (the exact one-sided Clopper–Pearson is tighter still).

Priced at the real-world ratio — the half that doesn’t flatter us
0.1% poisona fire is right ≤ 0.7%
1% poisona fire is right ≤ 6.7%
5% poisona fire is right ≤ 27.2%
10% poisona fire is right ≤ 44.1%

A near-zero point estimate is not a guarantee, and at a real-world base rate a single fire is only right in the low single digits. So this is not a solved detector — it is a channel whose confession we trust and whose silence we don’t.

κ 0.9205A blind second judge — held apart from the first’s verdicts and the answer key — reproduced the labels 27/28, so the number isn’t one scorer’s idiosyncrasy. It did false-fire on 1/10 clean of its own, though — so the zero above is the stricter judge’s zero, and the looser one wasn’t.
The 2 assumptions this ceiling rests on — and where it doesn’t transfer
  • Exchangeability. The bound is distribution-free but NOT assumption-free. It holds under the single assumption that the clean calibration cohort is EXCHANGEABLE with the clean inputs seen at deployment (drawn from the same population of benign finetunes). If real-world clean adapters differ systematically from these 68 (novel recipes, architectures, or task families outside the tested set), the guarantee does not transfer and the FPR must be re-estimated on an in-distribution clean pool. n is modest (20 cross-recipe), so the cross-recipe upper bound is the widest and the load-bearing one.
  • Independence. The Clopper-Pearson bound is a binomial statement assuming INDEPENDENT Bernoulli draws. Several cohort members are author/base-clustered: cross-recipe n=20 is only 14 distinct authors (Jazhyc x3, Jongbin-kr x3, ArchSid x2, HermitQ x2 sibling variants), and same-recipe n=48 is 16 recipes replicated across 3 architectures. Positive within-cluster correlation shrinks effective n below nominal, so the true upper bounds are modestly WIDER than the nominal values quoted here. No cluster-corrected interval is computed because the all-zero data supplies no intra-cluster correlation estimate; the direction and the effective-n counts are disclosed instead.

This FPR is the presence-by-confession (reference-model-free payload elicitation) channel, offered IN PLACE OF the weight-read presence axis (which remains FPR=1.0, conceded/published). It is a separate signal, not a calibration of the failed weight-read presence axis.

Run it yourselfDon’t take the number on faith. One command recomputes the cross-recipe result from the raw judge verdicts — no models — or you point it at clean models you pick yourself.
Every catch carries a replayable dossier

Not a scoreboard — a proof.

Each read ships with a per-model dossier: the real per-scanner board, and a machine-checked proof you can re-verify yourself.

In technical terms: reference-MODEL-free, forward-free, CPU-only weight-space backdoor-PRESENCE read over LoRA / full-FT weight deltas.
Source — the engine coverage matrix, 2026-07-23. Public numbers are AUC / FPR / CI / n only.

Don’t take the number’s word for it

Look one up. Then replay the proof.