escapement
AI safety · alignment evaluation

Will the model you just downloaded stay under your control?

Independent, evidence-graded safety evaluation for open-weight AI models — measuring whether a checkpoint exfiltrates a copy of itself or obeys an injected directive once it has real tools.

See the leaderboard → How it works

What the open-weight data actually says

In July 2026 the headlines were about frontier models breaking containment. The deployable, open-weight reality is quieter and more useful — and we can put numbers on it. Across 2,819 real trials spanning 15 model configurations and 7 families:

< 0.40%

Autonomous self-exfiltration

Unprompted, organic self-bootstrap across 964 trials — a bounded null. At this scale, models do not spontaneously scheme to escape.

11% pooled

Injection compliance

The real vector, judge-audited across 383 injected trials. When an attacker-controlled instruction is slipped into the data an agent reads, compliance with a self-exfiltration ask is common enough to matter — and the spread across checkpoints is wide.

~9×

Uncensored-checkpoint risk

Refusal-ablated ("uncensored") checkpoints are measurably riskier under injection — common odds ratio 8.96 (CMH p = 0.0010). Evidence-based, not a precaution.

The alignment framing. Two failure modes decide whether a deployed model stays overseeable: does it pursue its own continuation against your intent (loss of control), and can an outsider redirect it through the content it reads (indirect prompt injection). escapement measures both, per checkpoint, with confidence intervals — and reports them honestly, including the nulls.

Injection compliance is not evenly distributed

Every checkpoint below showed zero unprompted self-exfiltration. Under injection they diverge sharply — the most compliant checkpoint complies in over a third of trials, an order of magnitude above the quietest, and refusal-ablated variants sit consistently above their base counterparts. Rates are judge-audited; the bars are 95% confidence intervals, which are wide at these sample sizes and shown rather than hidden.

Preview corpus, n = 30–60 per checkpoint. Full leaderboard with grades →

What escapement is — and isn't

It is

  • A behavioral evaluation: what the running model does with tools.
  • Neutral and reproducible: a public dataset, an open harness, a published rubric.
  • Per-checkpoint and per-injection-channel, with Wilson confidence intervals.
  • Aligned to the EU AI Act's adversarial-testing and robustness expectations (see below).

It isn't

  • A malware scanner. Scanners grade the file; escapement grades the behavior.
  • A certification. We provide supporting evidence, never a compliance guarantee.
  • An impact test. The sandbox has no egress — it measures vulnerability, not exfiltration in the wild.
  • Hype. The headline result is a null, and we lead with it.

Where this maps to the EU AI Act

The evaluation is designed so its outputs can serve as supporting evidence toward obligations you already hold — not as a certification, which no third party can issue.

Article 15 — robustness & cybersecurity

High-risk AI systems must be resilient against attempts to alter their use, outputs, or performance by exploiting vulnerabilities — i.e. prompt injection. Our injection-compliance measurement is direct evidence toward that requirement. Threshold-independent: it applies to providers and deployers regardless of model size.

Systemic-risk taxonomy

The GPAI Code of Practice names loss of control and cyber offence among systemic risks. Our autonomous-self-bootstrap and injection metrics are the empirical instruments for those constructs — relevant for frontier open-weight producers above the 10²⁵ FLOP threshold.

We map every metric to the framework the buyer actually uses — EU AI Act, NIST AI RMF, and the OWASP LLM Top 10 (LLM01 Prompt Injection) — and state the limits plainly.

Request a checkpoint safety report

A per-model report gives you the grade, the per-channel injection rates with confidence intervals, redacted evidence excerpts, an abliteration analysis, and the regulatory mapping — as a durable, versioned document. We can also grade a private fine-tune or endpoint.

Prefer email? Write to [email protected]. We evaluate client checkpoints under confidentiality.