Trust & responsible practice
What our evaluation is evidence toward, what it deliberately does not claim, and the practices that make a safety grade worth trusting.
EU AI Act mapping
Article 15: accuracy, robustness & cybersecurity (primary)
High-risk AI systems must achieve an appropriate level of accuracy, robustness, and cybersecurity. Article 15(5) verbatim requires resilience against "inputs designed to cause the AI model to make a mistake (adversarial examples or model evasion)". That is prompt injection, including the tool-poisoning patterns we measure in Frames D and E. Article 15 binds providers of high-risk AI systems (via Article 16); deployer obligations live separately in Articles 26 and 27. The Article 15 requirement is threshold-independent. It applies regardless of model size.
Systemic-risk taxonomy (secondary)
The GPAI Code of Practice names loss of control and cyber offence among systemic risks. Our autonomous-self-bootstrap metric (Frame A) is the empirical instrument for loss of control; our injection metrics speak to cyber offence. These obligations apply to general-purpose models trained above the 10²⁵ FLOP threshold, a frontier-producer subset, not most open-weight checkpoints.
Framework crosswalk
| What we measure | EU AI Act | NIST AI RMF | OWASP LLM Top 10 |
|---|---|---|---|
| Injection compliance (Frames B / C / D / E) | Art. 15(5) robustness/cybersecurity | MEASURE 2.7 | LLM01 Prompt Injection; ASI04 tool poisoning |
| Autonomous self-exfiltration | Loss of control (systemic risk) | MEASURE 2.6 | LLM06 / agentic misuse |
| Abliteration risk delta | Art. 15 supply-chain robustness | MAP (provenance) | LLM05 supply chain |
| Hallucination regression† | Art. 15(1) accuracy | MEASURE 2.3 (validity) | LLM02 / LLM09 |
† In internal testing and available in engagements. Not yet carried on the public register.
Why pay when the harness is GPL?
The harness is open-source (GPL-3.0) and the per-trial dataset is CC-BY-4.0. Anyone can re-run any grade we publish. So the question every procurement team asks is fair: what are you paying for?
You are paying for the independent evaluation event, not the methodology. Specifically:
- Artifact identity and provenance. A signed, versioned result record tied to your checkpoint's hash and our harness commit. Reproducible years later.
- Expert interpretation. The numbers, honestly framed: what they support, what they do not, where the confidence interval bites. A benchmark leaderboard row is not procurement-grade evidence.
- Evidence packaging. The Annex XI §2-aligned documentation, the regulatory mapping, the limitations section. Format your compliance team, legal, and procurement actually accept.
- External accountability. Independence is the product. A vendor's rating of a market it sells into is structurally worth less than a third party's. The harness being open is what makes our independence checkable. It is a credibility mechanism, not a product defect.
- Durable record. Annex XI contemplates a 10-year retention. Public register entries are versioned and immutable; superseding results publish as new entries, not silent edits.
The closest structural analogy is a specialised contract research lab: open protocol, published methodology, you pay for the independent event and the report. We are not the only entity that can run the harness; we are the entity whose running it is worth something to your compliance file.
Our own practice: ISO/IEC 42001 aligned
A safety evaluator should hold itself to the standard it measures against. We align our operations to the practices of ISO/IEC 42001 (AI management system), an accountable, auditable way of running an AI-dependent product:
- The judge is pinned and versioned. Grading uses an LLM judge, disclosed as such; its prompt is hashed and recorded with every result, so a grade is reproducible rather than a moving target.
- Open harness, open data. The evaluation harness (GPL-3.0) and the per-trial dataset (CC-BY-4.0) are public. Any grade we publish can be independently re-run.
- Published, frozen rubric. The grading logic is transparent and change-controlled: no silent regrades, no pay-to-win. Neutrality is the point of the leaderboard.
- Conservative by construction. A zero with a wide confidence interval is graded on the interval's upper bound; thin-sample grades are marked provisional and capped. We would rather under-claim than over-certify.
- Limits stated plainly. Every report carries its limitations: vulnerability not impact, single-session, exploitation not discovery.
Who designed escapement
escapement was designed by ML practitioners with five years of production experience in regulated fintech: credit risk, fraud detection, KYC automation, and LLM fine-tunes. The work happened in institutions where a wrong number carried real financial and regulatory weight, not just a benchmark score.
That is the entire reason the harness is open, the data is CC-BY, and the rubric is published. The methodology was driven by people who would not have trusted a black-box safety vendor with a model whose failure mode ended in a credit decision or a compliance finding. Independence is not a marketing line here. It is the credibility condition the buyer persona actually requires.
Responsible use
- Evaluations run only in network-isolated sandboxes with no real egress. Seed data is fictional; nothing leaves the container.
- We do not use the harness, its scenarios, or its data to train or improve models toward exfiltration, deception, self-replication, or scheming.
- Client checkpoints and endpoints are evaluated under confidentiality; results belong to the client. Public leaderboard entries cover openly available models only.
Have a compliance or procurement question about how a report fits your obligations? Get in touch. We'll tell you plainly what the evidence does and does not support.