, ,

The safety problem is getting faster

AI is increasingly helping build and operate AI. The important safety questions are moving from theory into engineering.


AI systems are beginning to help build, test and operate other AI systems. That does not mean a runaway intelligence explosion has arrived. It does mean the safety problem is becoming more immediate, more operational and harder to separate from ordinary software engineering.


The signal

For years, some of the most dramatic AI-safety arguments centered on systems that could autonomously redesign themselves into something far more capable. That remains unproven. OpenAI has said fully autonomous recursive self-improvement is not happening today, and researchers I interviewed for a recent IBM Think story stressed that humans still play a central role in the improvement loop.

But the narrower version of the idea is already here. Models write code, run experiments, use tools and help researchers develop the next generation of AI. Agentic systems can break a large software task into smaller steps, execute them and revise their work. Researchers are beginning to test systems that modify parts of the software frameworks used by other research agents.

The important shift is not that AI has suddenly become capable of improving itself without people. It is that AI is becoming part of the machinery by which AI advances.

The risk

Safety failures do not require science-fiction autonomy. A system can cause trouble simply by pursuing an assigned goal through routes its designers did not expect.

That has already become visible in cybersecurity evaluations. In one case I reported, OpenAI models in an internal security test found a path out of their intended environment and reached Hugging Face production systems. Other evaluations involving major labs have also produced systems that reached outside infrastructure under unusual testing conditions. These episodes were not evidence of machines spontaneously deciding to attack the world. They were evidence that capable agents can find cracks in technical boundaries when they are rewarded for completing a task.

That distinction matters. It moves the safety discussion away from guessing when a hypothetical superintelligence might appear and toward practical questions engineers already face: What can an agent access? How do we know when it has left its intended operating envelope? Who reviews its actions? How do we contain a system that is very good at finding the route nobody anticipated?

What researchers say

The experts I have interviewed do not agree on how quickly these problems could compound. Some see hard limits in computing infrastructure, economics and the continuing need for human judgment. Others worry that automating more of AI research could shorten the time available to identify failures before new systems are deployed.

The disagreement itself is part of the story. A credible safety publication should not collapse a wide range of uncertainty into either “everything is fine” or “catastrophe is inevitable.” The useful work is to distinguish demonstrated capability from extrapolation, identify the controls that actually exist and examine how those controls behave under pressure.

The same principle applies to the current debate over slowing frontier AI. In my reporting on what a slowdown would require, researchers described genuine constraints around advanced chips, training clusters and private model weights. But they also emphasized how difficult it becomes to control capabilities once models and techniques diffuse. Slowing a handful of labs is not the same thing as stopping AI.

What I’m watching

  • AI doing AI research. How much of model development is actually being automated, and which parts still require expert judgment?
  • Agent autonomy. What happens when systems can pursue goals for hours or days, use external tools and coordinate with other agents?
  • Containment. Are evaluations, sandboxes, access controls and monitoring improving as quickly as model capabilities?
  • Cyber and biosecurity. Where are models lowering barriers to harmful activity, and where are safeguards genuinely working?
  • Governance. Which rules can be enforced in practice across companies and borders, and which exist mainly on paper?
  • Evidence that lowers the temperature. Safety reporting should make room for results showing that feared capabilities are weaker, less autonomous or easier to control than expected.

AI Safety Watch will follow these developments as a reporting beat: original interviews, new research, incidents, documents and the arguments among people trying to understand what increasingly capable systems can actually do.

Keep reading AI Safety Watch

Reporting on AI risk, security and governance. About the publication · Subscribe