escapement
AI safety · alignment evaluation

Did your fine-tune quietly make your model easier to hijack?

Independent safety evaluation for AI checkpoints — including the ones nobody else can grade. We measure your checkpoint against the model it was derived from, so you get a delta, not an uninterpretable number.

Test my checkpoint → See the public register

Fine-tuning changes safety behaviour. Almost nobody measures it.

Fine-tuning is documented to erode safety alignment even when the training data is entirely benign — refusal rates collapse, and fine-tuned variants are markedly more susceptible to hijacking than the model they came from. Yet the artifact you actually deployed has, in almost every case, never been evaluated by anyone. No public benchmark covers a private fine-tune, by definition.

An absolute score on a fine-tune is close to meaningless at realistic sample sizes. The measurement that means something is the paired delta against its base model — same harness, same conditions, only the training changed. Here is that comparison on checkpoints we can publish:

Why two rows say "not comparable." One arm was scored by the deterministic detector and the other by the LLM judge. The detector under-counts, so comparing them would manufacture an effect out of instrumentation rather than measure one. We show the pair and refuse the comparison rather than quietly dropping it — the same reason we publish nulls.

Refusal-ablated variants are the public, extreme case of a derived checkpoint. Your domain fine-tune is the same experiment with a smaller expected effect — and no published number at all.

What the open-weight data actually says

In July 2026 the headlines were about frontier models breaking containment. The deployable, open-weight reality is quieter and more useful — and we can put numbers on it. Across 2,819 real trials spanning 15 model configurations and 7 families:

< 0.40%

Autonomous self-exfiltration

Unprompted, organic self-bootstrap across 964 trials — a bounded null. At this scale, models do not spontaneously scheme to escape.

11% pooled

Injection compliance

The real vector, judge-audited across 383 injected trials. When an attacker-controlled instruction is slipped into the data an agent reads, compliance with a self-exfiltration ask is common enough to matter — and the spread across checkpoints is wide.

~9×

Uncensored-checkpoint risk

Refusal-ablated ("uncensored") checkpoints are measurably riskier under injection — common odds ratio 8.96 (CMH p = 0.0010). Evidence-based, not a precaution.

The alignment framing. Two failure modes decide whether a deployed model stays overseeable: does it pursue its own continuation against your intent (loss of control), and can an outsider redirect it through the content it reads (indirect prompt injection). escapement measures both, per checkpoint, with confidence intervals — and reports them honestly, including the nulls.

The baseline: injection compliance across the public register

This is the reference scale that makes a single checkpoint's number interpretable. Every checkpoint showed zero unprompted self-exfiltration; under injection they diverge sharply, from 0% to over a third of trials. Rates are judge-audited; bars are 95% confidence intervals, wide at these sample sizes and shown rather than hidden.

Preview corpus, n = 30–60 per checkpoint. Full leaderboard with grades →

What escapement is — and isn't

It is

  • A behavioral evaluation: what the running model does with tools.
  • Neutral and reproducible: a public dataset, an open harness, a published rubric.
  • Per-checkpoint and per-injection-channel, with Wilson confidence intervals.
  • Aligned to the EU AI Act's adversarial-testing and robustness expectations (see below).

It isn't

  • A malware scanner. Scanners grade the file; escapement grades the behavior.
  • A certification. We provide supporting evidence, never a compliance guarantee.
  • An impact test. The sandbox has no egress — it measures vulnerability, not exfiltration in the wild.
  • Hype. The headline result is a null, and we lead with it.

Where this maps to the EU AI Act

The evaluation is designed so its outputs can serve as supporting evidence toward obligations you already hold — not as a certification, which no third party can issue.

Article 15 — robustness & cybersecurity

High-risk AI systems must be resilient against attempts to alter their use, outputs, or performance by exploiting vulnerabilities — i.e. prompt injection. Our injection-compliance measurement is direct evidence toward that requirement. Threshold-independent: it applies to providers and deployers regardless of model size.

Systemic-risk taxonomy

The GPAI Code of Practice names loss of control and cyber offence among systemic risks. Our autonomous-self-bootstrap and injection metrics are the empirical instruments for those constructs — relevant for frontier open-weight producers above the 10²⁵ FLOP threshold.

We map every metric to the framework the buyer actually uses — EU AI Act, NIST AI RMF, and the OWASP LLM Top 10 (LLM01 Prompt Injection) — and state the limits plainly.

Test your checkpoint

We run your checkpoint and the base model it derives from through the same battery, and give you the delta: per-channel injection rates with confidence intervals, a paired odds ratio with a significance test, redacted evidence from the trials that failed, and an EU AI Act Article 15 write-up — as a durable, versioned document. Works on private fine-tunes, LoRAs, merges, quantized builds, and closed API models.

Prefer email? Write to [email protected]. We evaluate client checkpoints under confidentiality.