AI safety is no longer a single research community. It is a network of frontier labs, nonprofit research groups, evaluators, government institutes, policy organizations and individual researchers working on different versions of the same basic problem: how to understand, govern and control increasingly capable AI systems.
This is a living directory of the major people and organizations shaping that work. It is not meant to imply agreement among them. Some focus on catastrophic risk, some on misuse, some on model behavior and control, and some on governance or standards. The point is to make the field legible.
Frontier AI companies
Anthropic
Anthropic develops the Claude family of models and has made safety unusually central to its public identity. Its work includes mechanistic interpretability, model behavior, evaluations, frontier-risk frameworks and safeguards for increasingly capable systems. The company was founded by former OpenAI researchers and is structured as a public benefit corporation.
OpenAI
OpenAI develops frontier models and operates a formal Preparedness Framework for identifying and mitigating capabilities that could create severe harm. Its public safety work now spans cyber offense, chemical and biological risks, harmful manipulation, loss of control, safeguard circumvention and the governance of frontier systems.
Google DeepMind
Google DeepMind combines frontier-model development with dedicated technical safety, security and governance programs. Its Frontier Safety Framework addresses severe risks from advanced models, while its AGI Safety Council, led by co-founder Shane Legg, focuses on risks from more capable future systems.
Meta AI
Meta develops large-scale open and closed AI systems and publishes an Advanced AI Scaling Framework for assessing severe risks before deployment. Its framework covers areas including chemical and biological misuse, cybersecurity and loss of control, while Meta’s broader AI strategy remains unusually committed to open model releases.
xAI
xAI develops the Grok family of models. Its risk-management framework addresses malicious use, loss of control, security and third-party review, and the company publishes model cards and safety evaluations for frontier systems. Its safety approach is one of the newer institutional frameworks in the field.
Safe Superintelligence
Safe Superintelligence, or SSI, was founded by Ilya Sutskever and collaborators around a narrow mission: build superintelligent AI while making safety the central technical problem rather than a downstream deployment constraint. The company has kept a comparatively low public profile while recruiting researchers and raising substantial capital.
Technical safety and alignment research
LawZero
Founded by Yoshua Bengio in 2025, LawZero is pursuing what it calls safe-by-design AI. Its flagship research direction, Scientist AI, explores systems designed to make predictions without independently pursuing goals, an attempt to avoid some risks associated with increasingly agentic systems.
Alignment Research Center
The Alignment Research Center, or ARC, grew out of work by Paul Christiano and collaborators on scalable oversight and alignment. Its research has shifted over time, but the organization remains influential in attempts to understand how increasingly capable systems could be supervised when humans cannot directly verify every step of their work.
Redwood Research
Redwood Research is a nonprofit focused on AI safety and security. Its recent work is closely associated with AI control: how to safely use systems that might be capable of recognizing, exploiting or circumventing oversight. Buck Shlegeris is CEO and Ryan Greenblatt is chief scientist.
Apollo Research
Apollo Research studies deceptive behavior, scheming, loss of control and methods for monitoring advanced systems. Its work has helped turn questions once discussed largely in theoretical terms into empirical evaluation problems: whether models can strategically mislead evaluators, hide objectives or exploit weaknesses in oversight.
Machine Intelligence Research Institute
MIRI is one of the oldest organizations associated with technical AI alignment. It helped popularize many of the concepts that shaped early alignment research and now argues that current safety work is far behind frontier capabilities. Its public position is substantially more pessimistic than that of many other organizations in this directory.
FAR.AI
FAR.AI conducts technical safety research and runs programs intended to move ideas from papers into practical evaluation and engineering. Its work covers robustness, deception, adversarial behavior and methods for testing AI systems under difficult conditions. Adam Gleave is co-founder and CEO.
Center for Human-Compatible AI
UC Berkeley’s Center for Human-Compatible AI, or CHAI, studies how AI systems can remain compatible with human goals and preferences. The center is closely associated with Stuart Russell and with approaches that treat uncertainty about human objectives as a design feature rather than assuming those objectives can be perfectly specified in advance.
Center for AI Safety
The Center for AI Safety, led by Dan Hendrycks, conducts research, field-building and public education focused on large-scale risks from advanced AI. Its work spans robustness, evaluations, societal-scale hazards and the technical foundations of safer systems.
Independent evaluation and monitoring
METR
METR, short for Model Evaluation & Threat Research, independently evaluates frontier models and agents, particularly their ability to perform difficult tasks autonomously. The organization has worked with frontier labs and governments and has become a major reference point for measuring how quickly autonomous capabilities are improving. Beth Barnes is founder and CEO.
Palisade Research
Palisade Research studies dangerous or surprising behavior in advanced AI systems, including strategic behavior, cyber capabilities and attempts by models to resist or work around constraints. Its work sits between safety evaluation and adversarial security research.
Transluce
Transluce develops tools for understanding model behavior and internal representations. The organization, co-founded by Berkeley professor Jacob Steinhardt, is part of a broader push to make frontier models more legible to researchers rather than relying only on external behavior and benchmark scores.
Governance, standards and institutions
Centre for the Governance of AI
GovAI studies the institutions and policies that could shape advanced AI. Its researchers work on frontier-lab governance, international coordination, compute, standards and the economic and geopolitical consequences of increasingly capable systems. Ben Garfinkel is executive director.
Future of Life Institute
The Future of Life Institute works on the risks posed by powerful technologies and has become one of the most visible civil-society organizations in the AI-safety debate. Its work spans technical research, policy proposals, public campaigns and international governance. MIT professor Max Tegmark is founder and chair.
Frontier Model Forum
The Frontier Model Forum is an industry-supported nonprofit focused on frontier AI safety and security. Its members include Amazon, Anthropic, Google, Meta, Microsoft and OpenAI. The Forum develops shared practices, supports research and facilitates information sharing on risks such as cyber threats, CBRN misuse and advanced autonomous behavior.
UK AI Security Institute
The UK AI Security Institute is a government research organization that tests advanced systems, studies emerging risks and advises policymakers. Its work includes pre-deployment evaluations, model-security research and technical analysis of increasingly capable systems. Henry de Zoete is director and Jade Leung is chief technology officer.
U.S. Center for AI Standards and Innovation
CAISI, housed at the National Institute of Standards and Technology, develops voluntary standards and conducts evaluations of advanced AI systems. Its current work emphasizes demonstrable national-security risks, including cybersecurity, biosecurity and chemical threats, as well as standards for evaluating AI agents.
People to know
Yoshua Bengio
A Turing Award-winning pioneer of deep learning, Bengio is co-president and scientific director of LawZero, a professor at Université de Montréal and founder and scientific adviser to Mila. He has increasingly focused on catastrophic AI risk and on architectures intended to reduce the danger of highly agentic systems.
Geoffrey Hinton
One of the foundational figures in modern neural networks, Hinton is a University Professor Emeritus at the University of Toronto. Since leaving Google in 2023, he has become a prominent public voice warning that increasingly capable systems could become difficult to control and that researchers do not yet understand the full trajectory of the technology.
Stuart Russell
Russell is a UC Berkeley computer scientist, co-author of Artificial Intelligence: A Modern Approach and one of the field’s most influential advocates for building systems that remain uncertain about human objectives. He is closely associated with the Center for Human-Compatible AI.
Dario Amodei
Amodei is co-founder and CEO of Anthropic. A former OpenAI researcher, he helped build a company whose public identity rests heavily on the idea that frontier-model development and safety research need to advance together.
Daniela Amodei
Anthropic’s co-founder and president oversees the organization across research, engineering, product, governance and commercial operations. Her role puts her at the center of the institutional question facing frontier labs: how to translate a safety mission into the day-to-day incentives and decisions of a rapidly growing company.
Chris Olah
Olah is an Anthropic co-founder and the company’s interpretability research lead. His work helped establish mechanistic interpretability, the effort to understand the internal computations of neural networks rather than treating frontier models entirely as black boxes.
Amanda Askell
Askell is a philosopher and researcher at Anthropic whose work focuses on model behavior, values and alignment. She has played a central role in the development of Claude’s constitutional approach to specifying broad principles for model behavior.
Sam Altman
Altman is CEO of OpenAI and one of the most influential executives in frontier AI. His importance to AI safety is institutional as much as technical: deployment decisions, governance structures and the pace of model development at OpenAI all sit inside the broader argument over how quickly frontier capabilities should advance and under what safeguards.
Ilya Sutskever
Sutskever is a deep-learning pioneer, former OpenAI chief scientist and co-founder of Safe Superintelligence. He has long argued that superintelligent systems would require fundamentally stronger safety methods than current models and now leads a company built explicitly around that premise.
Demis Hassabis
Hassabis is co-founder and CEO of Google DeepMind. He oversees one of the world’s leading frontier AI laboratories, whose work now combines rapid capability advances with formal safety frameworks, internal review bodies and research on severe risks from more capable systems.
Shane Legg
Legg is a Google DeepMind co-founder and its chief AGI scientist. He leads the company’s AGI Safety Council, which focuses on extreme risks from highly capable future systems and advises on safety measures as frontier capabilities advance.
Rohin Shah
Shah is a Google DeepMind researcher whose work spans technical alignment, agent safety and control. He has been a prominent bridge between theoretical alignment arguments and empirical research on how increasingly autonomous systems should be evaluated and constrained.
Paul Christiano
Christiano founded the Alignment Research Center and has been one of the most influential technical thinkers in modern alignment research. His work has ranged from scalable oversight and reinforcement learning from human feedback to ways of supervising systems that may eventually outperform humans at important tasks.
Beth Barnes
Barnes is founder and CEO of METR. Her organization has become one of the field’s most important independent evaluators of frontier-model capabilities, especially the ability of AI agents to complete increasingly long and difficult tasks without human intervention.
Buck Shlegeris
Shlegeris is CEO of Redwood Research. His work is closely associated with AI control: how to use advanced systems safely when researchers cannot assume those systems will reliably cooperate with oversight.
Ryan Greenblatt
Greenblatt is Redwood Research’s chief scientist and a leading empirical researcher on AI control. His work tests whether potentially adversarial models can be used safely through monitoring, trusted components and other layers of defense.
Dan Hendrycks
Hendrycks is executive and research director of the Center for AI Safety. His work has covered benchmarks, robustness, model behavior and broader societal-scale risks from increasingly capable AI systems.
Adam Gleave
Gleave is co-founder and CEO of FAR.AI and a former DeepMind researcher. His work focuses on adversarial robustness, evaluations, deception and converting safety research into methods that can be tested under realistic pressure.
Geoffrey Irving
Irving is chief scientist at the UK AI Security Institute and has worked on scalable oversight, debate and methods for supervising systems that may be stronger than their human evaluators. His career spans both frontier labs and public-sector safety research.
Jan Leike
Leike is one of the best-known researchers associated with scalable alignment and superalignment. He previously led alignment work at OpenAI and later moved to Anthropic, where he has continued working on methods for supervising and controlling increasingly capable systems.
Max Tegmark
Tegmark is an MIT professor and founder and chair of the Future of Life Institute. His work spans technical safety, public advocacy and policy, and he has become one of the most visible proponents of stronger safeguards around advanced AI.
Ben Garfinkel
Garfinkel is executive director of the Centre for the Governance of AI. His work sits at the intersection of technical progress and institutional response, with an emphasis on how governments, labs and international bodies might govern increasingly powerful systems.
Jade Leung
Leung is chief technology officer of the UK AI Security Institute and also serves as the British prime minister’s AI adviser. She previously led governance work at OpenAI and now sits at the center of the UK government’s technical response to frontier-model risks.
Chris Meserole
Meserole is executive director of the Frontier Model Forum. He leads an industry-backed effort to develop shared safety practices, improve information sharing and coordinate research around frontier risks.
About this directory: AI Safety Watch updates this page as organizations form, merge, change leadership or move into new areas of research. Inclusion does not imply endorsement. If an important organization or researcher is missing, send a note.




