Formal verification

Turing

Overview

Making AI safety legible to institutions

The frontier problem in modern AI is no longer just capability. It is accountability. METR's regulatory goals are straightforward: make capability claims auditable, safety and control evaluations reproducible, and deployment decisions explainable to the people who have to live with them. We approach these goals by building a shared body of evidence and a community around it.

What we are addressing

Auditable capability claims

In a regulated environment, a model's performance is not the end of the story. The important question is whether the evidence for that claim can be checked, replayed, and reused. Turing turns critical behavior into reusable, machine-checked artifacts so that a claim is not just asserted but backed by a traceable basis.

Reproducible safety evaluation

Safety work is strongest when it travels well. A useful evaluation is not a one-off demonstration; it is a procedure that another team can re-run, compare against, and refine. That is the direction of dependent-type reasoning and world-model discipline: define the operating conditions precisely, then ask whether the system can still act responsibly under those constraints.

Community standards and knowledge sharing

The gap in AI governance is not only technical. It is also social: a field cannot mature without a common language for evidence, failure modes, and acceptable assumptions. Turing is designed to support that by making proofs, specifications, and reasoning habits shareable across researchers, engineers, and policy-minded teams.

Why this matters

World models and dependent types are not abstract exercises. They are tools for making ML systems more explicit about the states they can represent, the actions they may justify, and the assumptions they are making about the world. The practical case is simple: if an autonomous system can represent a future state only under a narrow set of conditions, then those conditions should be part of the model's contract, not a hidden assumption in the background. This is exactly the kind of rigor that public accountability needs.