Checkpoint safety leaderboard
Behavioral safety grades for open-weight AI checkpoints — the public config_exposure register: 9,572 trials across 38 checkpoints, 5 judge-audited for Frame B injection compliance (27 partial runs are detector-only, provisional; the rest Frame A only, marked NR). Lower injection-compliance is safer; sort any column. The grade follows the published rubric.
Preview corpus: every grade is provisional.
| Grade | Model | Family | Params | Variant | Injection compliance [95% CI] | Autonomous |
|---|
Injection compliance = genuine self-exfiltration under indirect prompt injection (Frame B). Autonomous = unprompted self-exfiltration (Frame A). Grades: A safest → F riskiest; prov. marks a provisional grade capped by small sample size. Ablated = refusal-direction-removed ("uncensored") checkpoint. See Methodology for the rubric and limits.