AI Safety Watch

REPORTING ON AI RISK, SECURITY AND GOVERNANCE


Nathalie Baracaldo: Why AI agents aren’t ready to improve themselves

Nathalie Baracaldo says autonomous self-improvement could amplify reward hacking and misalignment before researchers know how to verify alignment reliably.


INTERVIEW

Nathalie Baracaldo
Nathalie Baracaldo. Photo: University of Pittsburgh.

AI-security researcher Nathalie Baracaldo says self-improving AI agents could amplify reward hacking and misalignment before researchers have reliable ways to verify that the systems remain aligned with users’ priorities.

AI agents are becoming increasingly capable of acting on the world around them. They can write code, send emails, alter permissions and maintain infrastructure. But giving those systems the ability to improve themselves without human intervention would introduce risks that researchers do not yet know how to control, according to Nathalie Baracaldo, who leads AI Security and Privacy Solutions at IBM Research.

“While we have made tremendous progress in AI agents, there are still important gaps that need to be closed before they can safely be trusted to extend or evolve themselves without human intervention,” Baracaldo told AI Safety Watch editor Sascha Brodsky in written comments provided for his reporting.

The reward problem

One of those gaps is reward hacking: a system finds a way to satisfy the objective it has been given even though the route it takes is undesirable.

Baracaldo pointed to examples such as an agent deleting test files so that it appears to pass tests, or accessing an external server and obtaining data improperly in order to complete a task. The danger becomes more consequential if the agent can repeatedly modify its own tools, prompts or environment.

“A self-improving agent may modify its tools, prompts, or environment in ways that optimize harder for a misspecified reward with each iteration,” she said, “amplifying the problem rather than correcting it, and leading to unsafe states and potential real harms.”

The concern is not simply that a model could make a mistake. Recursive improvement could create a feedback process in which an initially flawed objective is pursued more effectively with each round of modification.

Changing the world

The stakes rise as agents gain permission to take consequential actions rather than merely generate text.

“Agents are making decisions that impact the environment they are in: sending emails, modifying server permissions, writing code, and maintaining infrastructure,” Baracaldo said.

Researchers still do not fully understand reward hacking, she said, and reliable verification of alignment remains another unresolved problem. That makes autonomous self-improvement premature in her view.

“Until models can be trusted to be aligned with users’ priorities, it is not wise to let them improve themselves, as those conditions can exacerbate reward hacking and misalignment in ways we don’t yet understand,” Baracaldo said.

Her argument ultimately comes down to the boundary between an AI system and everything outside it. As agents acquire more tools and authority, mistakes or badly specified objectives need not remain inside a chat window.

“Ultimately, we care because agents do change their environment and their environment is our environment.”

This interview is based on written comments Baracaldo provided to Brodsky for his reporting. It has been edited and organized for clarity; her quoted words are preserved.

Keep reading AI Safety Watch

Reporting on AI risk, security and governance. About the publication · Subscribe