AI Safety Watch

REPORTING ON AI RISK, SECURITY AND GOVERNANCE


Why an AI explanation may sound right and still be wrong

David Cox says fluent self-explanations can tempt people to anthropomorphize AI. Mechanistic interpretability aims at the harder problem: understanding what is actually happening inside a model.


INTERVIEW

When an AI system explains why it did something, the explanation may be coherent, detailed and completely plausible. That does not make it true.

David Cox sees that gap as one reason researchers need interpretability techniques that look beyond a model’s own account of its behavior.

“A perfectly valid sounding explanation” can still be false, Cox said. The distinction, he argued, is between asking a model to generate a convincing story about itself and developing methods that let researchers inspect the mechanisms that produced its behavior.

That problem is easy to blur because large language models speak so fluently. People naturally reach for human terms such as understanding, attention, introspection and deception. Cox warned that those words can smuggle assumptions into the discussion.

“The words we use shape how we think about things,” he said. Terms borrowed from psychology or neuroscience may have a different technical meaning in AI, even when ordinary language makes the concepts sound equivalent.

Cox is especially cautious about claims that a model is conscious or possesses humanlike self-knowledge. He said those claims are difficult to turn into falsifiable scientific hypotheses. A system may be able to monitor some of its internal state because doing so helps it perform a task, without that capability implying humanlike introspection.

The safety problem is nevertheless real. Models can produce behavior that, if a human produced it, might be described as deceptive. Cox’s objection is to assuming a motive rather than studying the mechanism.

That is where mechanistic interpretability becomes important. Instead of relying only on what a model says about itself, researchers try to identify what information the model represents and how those internal representations contribute to behavior.

Cox said the purpose is practical: researchers want systems that can be controlled well enough to do the intended task without producing unwanted behavior. Better probes into a model’s internals could help distinguish genuine mechanisms from persuasive self-description.

He also urged caution when researchers deliberately put agents into role-playing scenarios designed to elicit extreme behavior. Such experiments can reveal vulnerabilities, but they can also tempt observers to read human motives into software.

“We want to be able to control them,” Cox said. “We want to be able to have them do what we need them to do and not do things that aren’t what we need them to do.”

This interview was originally conducted by Brodsky for IBM Think. Material that Cox indicated required separate clearance has been omitted.

Keep reading AI Safety Watch

Reporting on AI risk, security and governance. About the publication · Subscribe