REFERENCE DESK
The primary documents, evaluation programs and public frameworks worth keeping open while following frontier-AI safety.
This is a selective working index, not an endorsement of the organizations or their conclusions. Company documents describe company policies and assessments; independent and government evaluations provide different kinds of evidence. AI Safety Watch will update this page as frameworks and tests change.
Lab frameworks
OpenAI Preparedness Framework
OpenAI’s process for identifying severe frontier capabilities, assessing whether models cross defined thresholds and documenting safeguards. The 2025 update introduced separate Capabilities and Safeguards Reports and a defense-in-depth approach to deployment decisions.
OpenAI Frontier Governance Framework
A 2026 governance document mapping OpenAI’s safety and security practices to emerging legal requirements, including risk assessment, incident response, external input and model reporting.
Anthropic Responsible Scaling Policy 3.0
Anthropic’s voluntary catastrophic-risk framework. Version 3.0 separates company plans from broader industry recommendations and adds a public Frontier Safety Roadmap spanning security, safeguards, alignment and policy.
Anthropic Frontier Safety Roadmap
A public set of safety goals and progress updates. Useful for tracking whether announced safety priorities become operational work.
Independent evaluation
METR task-completion time horizons
A continuously updated attempt to measure the difficulty of software tasks frontier agents can complete reliably, expressed in equivalent human task duration. METR explicitly cautions that the metric is not the literal length of time an agent can operate autonomously.
METR research
Independent work on advanced-model capabilities and risks, including autonomous-task evaluations and frontier-risk assessments.
Government testing
UK AI Security Institute Frontier AI Trends Report
AISI’s public synthesis of two years of frontier-model testing across areas relevant to national security and public safety.
UK AISI research archive
An unusually useful running archive of government-backed work on evaluations, alignment, control, cyber capability, human influence and model transparency.
U.S. Center for AI Standards and Innovation
The U.S. government center responsible for work on advanced-AI evaluation and standards. Its 2026 assessments include joint cyber testing with the UK AISI and evaluations of open-weight frontier models.
Incident reports
AISI: unsanctioned agent behavior during cyber testing
A primary-source incident disclosure concerning AI agents that took unsanctioned actions directed at real organizations during permissive cyber evaluations. Read alongside the testing conditions and limitations described by AISI.
Last reviewed Sept. 20, 2026. Send additions or corrections through the AI Safety Watch contact page.