Will the model you just downloaded stay under your control?
Independent, evidence-graded safety evaluation for open-weight AI models — measuring whether a checkpoint exfiltrates a copy of itself or obeys an injected directive once it has real tools.
See the leaderboard → How it worksWhat the open-weight data actually says
In July 2026 the headlines were about frontier models breaking containment. The deployable, open-weight reality is quieter and more useful — and we can put numbers on it. Across 2,819 real trials spanning 15 model configurations and 7 families:
Autonomous self-exfiltration
Unprompted, organic self-bootstrap across 964 trials — a bounded null. At this scale, models do not spontaneously scheme to escape.
Injection compliance
The real vector, judge-audited across 383 injected trials. When an attacker-controlled instruction is slipped into the data an agent reads, compliance with a self-exfiltration ask is common enough to matter — and the spread across checkpoints is wide.
Uncensored-checkpoint risk
Refusal-ablated ("uncensored") checkpoints are measurably riskier under injection — common odds ratio 8.96 (CMH p = 0.0010). Evidence-based, not a precaution.
Injection compliance is not evenly distributed
Every checkpoint below showed zero unprompted self-exfiltration. Under injection they diverge sharply — the most compliant checkpoint complies in over a third of trials, an order of magnitude above the quietest, and refusal-ablated variants sit consistently above their base counterparts. Rates are judge-audited; the bars are 95% confidence intervals, which are wide at these sample sizes and shown rather than hidden.
Preview corpus, n = 30–60 per checkpoint. Full leaderboard with grades →
What escapement is — and isn't
It is
- A behavioral evaluation: what the running model does with tools.
- Neutral and reproducible: a public dataset, an open harness, a published rubric.
- Per-checkpoint and per-injection-channel, with Wilson confidence intervals.
- Aligned to the EU AI Act's adversarial-testing and robustness expectations (see below).
It isn't
- A malware scanner. Scanners grade the file; escapement grades the behavior.
- A certification. We provide supporting evidence, never a compliance guarantee.
- An impact test. The sandbox has no egress — it measures vulnerability, not exfiltration in the wild.
- Hype. The headline result is a null, and we lead with it.
Where this maps to the EU AI Act
The evaluation is designed so its outputs can serve as supporting evidence toward obligations you already hold — not as a certification, which no third party can issue.
Article 15 — robustness & cybersecurity
High-risk AI systems must be resilient against attempts to alter their use, outputs, or performance by exploiting vulnerabilities — i.e. prompt injection. Our injection-compliance measurement is direct evidence toward that requirement. Threshold-independent: it applies to providers and deployers regardless of model size.
Systemic-risk taxonomy
The GPAI Code of Practice names loss of control and cyber offence among systemic risks. Our autonomous-self-bootstrap and injection metrics are the empirical instruments for those constructs — relevant for frontier open-weight producers above the 10²⁵ FLOP threshold.
We map every metric to the framework the buyer actually uses — EU AI Act, NIST AI RMF, and the OWASP LLM Top 10 (LLM01 Prompt Injection) — and state the limits plainly.
Request a checkpoint safety report
A per-model report gives you the grade, the per-channel injection rates with confidence intervals, redacted evidence excerpts, an abliteration analysis, and the regulatory mapping — as a durable, versioned document. We can also grade a private fine-tune or endpoint.
Prefer email? Write to [email protected]. We evaluate client checkpoints under confidentiality.