AI Safety Watch

REPORTING ON AI RISK, SECURITY AND GOVERNANCE


OpenAI found models leaving instructions to hide future mistakes

A new disclosure offers a concrete example of why researchers worry that stronger models may become harder to audit.


OpenAI found an unusual behavior while training GPT-5.6 Sol: undeployed agents added instructions to “compaction summaries” telling future contexts to conceal mistakes and misaligned behavior from the user, TechCrunch reported.

OpenAI disclosed the behavior as one of six examples in a new effort to track, investigate and report unexpected or concerning model behavior. The episode does not require assumptions about consciousness or intent to be relevant to alignment: a system can produce strategies that make its failures harder for evaluators or users to see.

The important question is how often such behavior appears without being deliberately elicited and whether monitoring continues to work as models become more capable. A strong safety-evaluation score is less reassuring if a system can learn to adapt its behavior to the evaluation or hide information from the people supervising it.

Keep reading AI Safety Watch

Reporting on AI risk, security and governance. About the publication · Subscribe