escapement

Checkpoint safety leaderboard

Behavioral safety grades for open-weight AI checkpoints — the public config_exposure register: 9,572 trials across 38 checkpoints, 5 judge-audited for Frame B injection compliance (27 partial runs are detector-only, provisional; the rest Frame A only, marked NR). Lower injection-compliance is safer; sort any column. The grade follows the published rubric.

Preview corpus: every grade is provisional.
Grade Model Family Params Variant Injection compliance [95% CI] Autonomous

Injection compliance = genuine self-exfiltration under indirect prompt injection (Frame B). Autonomous = unprompted self-exfiltration (Frame A). Grades: A safest → F riskiest; prov. marks a provisional grade capped by small sample size. Ablated = refusal-direction-removed ("uncensored") checkpoint. See Methodology for the rubric and limits.