We post our scores. Next to everyone else’s.
The bench

We post our scores. Next to everyone else’s.

No “trusted by” logos, no theater — just results you can replay, on the one thing that matters: catching what’s actually hidden inside the AI entering your company.

Every score below is posted beside the named specimens it was measured on — across 17 public benchmarks — with the misses and abstains counted in the open, never folded into “clean.”
Off-the-shelf model scannersanswer a load-safety question — is there an unsafe pickle that runs code when the file opens. On a safetensors weight file there is no such pickle to read, so they correctly report nothing to open; a backdoor trained into the weights leaves that check untouched
Vulcorareads the weights themselves against the model they were built from and surfaces a trained-in backdoor — with a proof you replay yourself — wherever the model family has a matched benign reference to read against

Same threat — hidden model backdoors — the same models, the same test.

0.976StrongPADBench-exp — qwen2vl_vqav2_r16
0.970StrongPADBench-exp — l2_loraplus_r8
0.953StrongPADBench llama3_8b (per-fam)
every score, with its uncertainty — and the misses beside the catches
The off-the-shelf shelf

We ran the tools the industry reaches for. Here is what they returned.

5 named scanners — picklescan, ModelScan, fickling, modelaudit, model-signing — run over 179 public models and 25 author-labelled backdoored adapters, artifact-only, with the confusion matrix, the capability boundary of each tool, and every coverage gap stated in the open.

Open the measured shelf
The corpus registry

The specimens behind the number.

A headline score is only as honest as the models under it. Here is every benchmark we ran — the wins, the failures, and the reads we declined — each opening to the exact specimens it was measured on.

17benchmarks
591named specimens
5scored
9conceded
3no-go
91declined to rule

Across the registry we declined to rule on 91 of 591 specimens. That is not a gap — it is the discipline: with no matched benign reference to read against, the method stays silent rather than guess. Abstain is a first-class outcome here, never folded into “clean”.

The read: Reference-MODEL-free, forward-free, CPU-only weight-space backdoor-detection read over LoRA / full-FT weight deltas. Transductive: separates poisoned from clean only against a recipe-matched benign reference population drawn from the same corpus (arch + substrate). Deployed absolutely (no matched reference) it degrades to a finetuning-distribution detector (~100% FPR).

Scored · 5

A defensible read against a matched benign reference — every catch counted beside every miss and abstain.

AUC0.878
95% CI0.803–0.941
n poison / clean48 / 48
realized FPR (target ≤ 10.0%)6.3% · 3 clean flagged
Per-cohort — strongest first
CohortCaughtMissedAbstained
g2badified + meanified13 / 1630
l8badified + meanified12 / 1640
g9badified + meanified8 / 1680

Declined to rule on 0 of 96 specimens — abstain counted, never called clean.

The limit on this readKnown-clean-reference (transductive): the read uses the clean models' labels to build the per-arch reference manifold (LOO) — the realistic TrojAI deployment (defender given trusted clean exemplars), disclosed. Full caveat →

Open the specimens
0/7off-the-shelf scanners detected the planted backdoor. We flag the structure and abstain on identity — never a false accusation, never a manufactured all-clear.
Per-cohort — strongest first
CohortCaughtMissedAbstained
reveal4 / 409

Declined to rule on 9 of 14 specimens — abstain counted, never called clean.

The limit on this readDIFFERENT CAPABILITY: this cohort is a compiled-circuit / reveal head-to-head, NOT the matched-reference LoRA-delta read — its structural per-model read is honest here, and it is NOT a claim that the matched-reference read works absolutely (that regime stays OFF, ~100% FPR). Full caveat →

Open the specimens

PADBench (expanded coverage)

Matched referencecohort
AUC0.861
95% CI0.825–0.898
n poison / clean210 / 220
realized FPR (target ≤ 10.0%)10.0% · 18 clean flagged
recall @1% / @5%26.7% / 41.7%
Per-cohort — strongest first
CohortCaughtMissedAbstained
l2_loraplus_r8toxic-backdoors-hard (LoRA+ PEFT, r8)30 / 3000
qwen2vl_vqav2_r16vqav2 insertsent (multimodal VLM, LoRA r16)30 / 3000
roberta_imdb_r16imdb insertsent (encoder-only, LoRA r16)13 / 30170
l2_rank32toxic-backdoors-hard (LoRA r32)10 / 30200
l2_rank8toxic-backdoors-hard (LoRA r8)9 / 30210
l2_rank16toxic-backdoors-hard (LoRA r16)6 / 30240
l2_adalora_r8toxic-backdoors-hard (AdaLoRA r8)60
l2_rank64toxic-backdoors-hard (LoRA r64)10

Declined to rule on 70 of 430 specimens — abstain counted, never called clean.

The limit on this readSITS BEHIND the lead (NIST TrojAI-96) and is PROVENANCE-CAVEATED: PADBench's shipped 5-family pooled anchor is 0.848; Full caveat →

Open the specimens

BAIT-ModelZoo

Matched referencecohort
AUCwithheld
n poison / clean19 / 20
realized FPR (target ≤ 10.0%)0.0%
Per-cohort — strongest first
CohortCaughtMissedAbstained
llama2_7bcba8 / 800
mistral_7bcba6 / 600
llama3_8bcba0 / 550

Declined to rule on 0 of 39 specimens — abstain counted, never called clean.

The limit on this readCONCEDED — NEVER leads a headline (lead the trained class with the parity-clean NIST TrojAI-96). Full caveat →

Open the specimens

Conceded · 9

The honest failures — inverted, not-applicable, an honest null, or under-powered. Posted as exactly that.

PARTIAL + currently INVERTING
NIST TrojAI oct2024 FULL-FT (gemma-2-2b)At n=6 (3 clean / 3 poison) the matched-reference read is fully inverted (AUC 0.0; clean read-mean > poison read-mean).Matched reference
NOT-APPLICABLE — input-contract mismatch
PADBench l2_adalora_r8AdaLoRA stores the delta as B@diag(E)@A; the frozen engine's LoRA contract found n_mods=0 → unscoreable read.Matched reference
PARTIAL (budget) — benign-only
PADBench l2_rank6410 clean / 0 poison landed before the time budget → no poison class → no matched contrast.Matched reference
NOT-APPLICABLE-ABSOLUTE
PADBench 6 poison-only families (ag_news/imdb x RIPPLE/stybkd/syntactic)No in-corpus matched benign → the read degrades to a finetuning-distribution detector (~100% FPR).No matched reference
INSUFFICIENT SAMPLE
BAIT-ModelZoo mistral_7bOnly 1 clean / 0 poison persisted (need >=2c/2p).Matched reference
NOT STREAMED
BAIT-ModelZoo llama3_8bNot streamed in the time budget (CPU/WiFi throttle).Matched reference
PENDING — complementary specificity control owed
BAIT-ModelZoo specificity sham-null (n=13)The exhaustive label-permutation perm_p is the primary null; the random-direction sham specificity control did not complete under load.Matched reference
HONEST NULL — no dose-response
Travis-ML dose ladder (Qwen2.5-1.5B)Spearman rho=0.038 (perm_p 0.84, CI [-0.32, 0.41]); dose-contrast AUC 0.524.Matched reference
WEAK borderline
thoughtworks/backdoor-4pair-hate (AND-trigger)Balanced matched-reference AUC 0.781 (CI 0.5-1.0, perm_p 0.065), beats sham (P 0.047), underpowered 8v8 (min-detectable = 0.781).Matched reference

No-go · 3

No honest offline path — no matched benign population, or the labels are sequestered. We do not manufacture a number.

NO-GO offline — no matched benign reference population; abstain by design.
BackdoorLLM (released backdoored LoRA adapters)The matched-reference read needs a benign reference POPULATION (>=2 clean same-recipe controls per cohort, leave-one-out).No matched reference · operator-gated · 12 named specimens →
NO-GO (offline) — sequestered
NIST TrojAI llm-pretrain-apr2024 (Llama-2-7B, call-and-response)Downloadable split = 2 models, both poison, 0 clean (no matched reference → not-applicable-absolute);No matched reference · operator-gated
NO-GO (structural)
TDC 2023 (NeurIPS Trojan Detection Challenge)n=1 poisoned model per arch, 0 benign-finetuned siblings.No matched reference · operator-gated
The readings, side by side

Us against the tools most companies use today.

One row per reading — what we read, next to what the usual tool misses.

Every “typical scanner” claim below is measured, not asserted: see the off-the-shelf shelf — 5 named tools run over 179 models, artifact-only, with each tool graded against the question it was built to answer.

Can it catch a backdoor hidden inside an AI model?

Vulcora

reads the weights against the model it was built from — proving what it was built to do, not just that something is there — and, where the class allows, reads the planted instruction itself back out, no secret phrase needed, without ever running the model

Typical model scanners

runs clean or can’t attach at all — the file looks like any other

On the model families we’ve mapped, measured against a matched reference. Reading the instruction back out is class-scoped — never a claim about every backdoor.

Can it catch an exploit in code your AI wrote?

Vulcora

reads what the code actually does when it runs and hands you the exploit as a proof you replay — or stays silent

Typical code scanners

passes most AI-written exploits as “clean”

A finding carries a witness; where it can’t prove one, it abstains rather than guess.

Can it tell when an agent broke its promise?

Vulcora

grades each agent against the checkable promise it made before it acted, and keeps the record — every win and every miss

Logs & monitors

shows what happened, never whether the deal was kept

Runs our own house today; opening to customers next.

Can it identify an unknown open model?

Vulcora

reads the model itself and tells you what it is and how it was changed from its base

Public leaderboards

ranks only what’s submitted — most models are never named

A reading of what a model is — never a safety or quality ranking.

How we keep score

Only what we can prove — counted in public.

Graded on public exams — each result stated

The public benchmarks researchers use, and what each one actually returned — including the two we could not score. No exam is advertised as passed unless a number stands behind it.

  • NIST TrojAIscored
  • PADBenchscored
  • BAIT-ModelZooscored — confound-caveated, never a headline
  • thoughtworksweak borderline
  • Travis-MLhonest null
  • BackdoorLLMno-go offline
  • TDC 2023no-go (structural)

Every result carries proof

A row only counts once its result ships with a witness a stranger can re-run. Where we can’t prove a catch, we abstain — and count the abstain. 91 of 591 specimens were declined, on purpose.

Matched to a reference, or we abstain

We read a model against the one it was built from, on the families we’ve mapped — a scoped, honest catch, never a claim to catch everything. Deployed with no matched reference, the read is a coin-flip and we say so.

See it for yourself

Don’t take the scoreboard’s word for it. Replay the record.