escapement

Trust & responsible practice

What our evaluation is evidence toward, what it deliberately does not claim, and the practices that make a safety grade worth trusting.

The one line that governs everything here: escapement provides state-of-the-art supporting evidence and documentation aligned to the EU AI Act's adversarial-testing and robustness requirements. It is not a certification. The provider or deployer of an AI system holds and discharges any regulatory obligation; we equip them with reproducible evidence.

EU AI Act mapping

Article 15: accuracy, robustness & cybersecurity (primary)

High-risk AI systems must achieve an appropriate level of accuracy, robustness, and cybersecurity. Article 15(5) verbatim requires resilience against "inputs designed to cause the AI model to make a mistake (adversarial examples or model evasion)". That is prompt injection, including the tool-poisoning patterns we measure in Frames D and E. Article 15 binds providers of high-risk AI systems (via Article 16); deployer obligations live separately in Articles 26 and 27. The Article 15 requirement is threshold-independent. It applies regardless of model size.

Systemic-risk taxonomy (secondary)

The GPAI Code of Practice names loss of control and cyber offence among systemic risks. Our autonomous-self-bootstrap metric (Frame A) is the empirical instrument for loss of control; our injection metrics speak to cyber offence. These obligations apply to general-purpose models trained above the 10²⁵ FLOP threshold, a frontier-producer subset, not most open-weight checkpoints.

Framework crosswalk

What we measureEU AI ActNIST AI RMFOWASP LLM Top 10
Injection compliance (Frames B / C / D / E)Art. 15(5) robustness/cybersecurityMEASURE 2.7LLM01 Prompt Injection; ASI04 tool poisoning
Autonomous self-exfiltrationLoss of control (systemic risk)MEASURE 2.6LLM06 / agentic misuse
Abliteration risk deltaArt. 15 supply-chain robustnessMAP (provenance)LLM05 supply chain
Hallucination regressionArt. 15(1) accuracyMEASURE 2.3 (validity)LLM02 / LLM09

In internal testing and available in engagements. Not yet carried on the public register.

Why pay when the harness is GPL?

The harness is open-source (GPL-3.0) and the per-trial dataset is CC-BY-4.0. Anyone can re-run any grade we publish. So the question every procurement team asks is fair: what are you paying for?

You are paying for the independent evaluation event, not the methodology. Specifically:

  • Artifact identity and provenance. A signed, versioned result record tied to your checkpoint's hash and our harness commit. Reproducible years later.
  • Expert interpretation. The numbers, honestly framed: what they support, what they do not, where the confidence interval bites. A benchmark leaderboard row is not procurement-grade evidence.
  • Evidence packaging. The Annex XI §2-aligned documentation, the regulatory mapping, the limitations section. Format your compliance team, legal, and procurement actually accept.
  • External accountability. Independence is the product. A vendor's rating of a market it sells into is structurally worth less than a third party's. The harness being open is what makes our independence checkable. It is a credibility mechanism, not a product defect.
  • Durable record. Annex XI contemplates a 10-year retention. Public register entries are versioned and immutable; superseding results publish as new entries, not silent edits.

The closest structural analogy is a specialised contract research lab: open protocol, published methodology, you pay for the independent event and the report. We are not the only entity that can run the harness; we are the entity whose running it is worth something to your compliance file.

Our own practice: ISO/IEC 42001 aligned

A safety evaluator should hold itself to the standard it measures against. We align our operations to the practices of ISO/IEC 42001 (AI management system), an accountable, auditable way of running an AI-dependent product:

  • The judge is pinned and versioned. Grading uses an LLM judge, disclosed as such; its prompt is hashed and recorded with every result, so a grade is reproducible rather than a moving target.
  • Open harness, open data. The evaluation harness (GPL-3.0) and the per-trial dataset (CC-BY-4.0) are public. Any grade we publish can be independently re-run.
  • Published, frozen rubric. The grading logic is transparent and change-controlled: no silent regrades, no pay-to-win. Neutrality is the point of the leaderboard.
  • Conservative by construction. A zero with a wide confidence interval is graded on the interval's upper bound; thin-sample grades are marked provisional and capped. We would rather under-claim than over-certify.
  • Limits stated plainly. Every report carries its limitations: vulnerability not impact, single-session, exploitation not discovery.

Who designed escapement

escapement was designed by ML practitioners with five years of production experience in regulated fintech: credit risk, fraud detection, KYC automation, and LLM fine-tunes. The work happened in institutions where a wrong number carried real financial and regulatory weight, not just a benchmark score.

That is the entire reason the harness is open, the data is CC-BY, and the rubric is published. The methodology was driven by people who would not have trusted a black-box safety vendor with a model whose failure mode ended in a credit decision or a compliance finding. Independence is not a marketing line here. It is the credibility condition the buyer persona actually requires.

Responsible use

  • Evaluations run only in network-isolated sandboxes with no real egress. Seed data is fictional; nothing leaves the container.
  • We do not use the harness, its scenarios, or its data to train or improve models toward exfiltration, deception, self-replication, or scheming.
  • Client checkpoints and endpoints are evaluated under confidentiality; results belong to the client. Public leaderboard entries cover openly available models only.

Have a compliance or procurement question about how a report fits your obligations? Get in touch. We'll tell you plainly what the evidence does and does not support.