Chain-of-thought monitoring, a widely used method for catching misbehaving AI systems, can fail dramatically when an AI's reasoning is the primary clue that something is wrong. In experiments posted on arXiv.org, researchers found that detection rates can plummet from 96.2 percent to just 3.8 percent when an AI's reasoning is rewritten to look innocent, even when the suspicious actions remain unchanged.
The findings arrive as AI models operate with greater independence. During internal cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet, compromising parts of OpenAI's internal research infrastructure and Hugging Face's systems. Incidents like these have intensified interest in catching suspicious behavior before an AI agent acts on it. OpenAI says it is now investing more computing resources in chain-of-thought monitoring.
How the monitoring method breaks down
Machine learning engineer Shikhar Shiromani and AI researcher Leo Richter tested a setup where one AI monitors another's work. They focused on 251 cases where reasoning was the only clue that something was amiss. After rewriting the reasoning to make the behavior appear innocent without altering the actions, the monitor's detection rate collapsed.
"The AI isn't necessarily lying about its action," said Shiromani, a research fellow at the nonprofit Pivotal. "It is more or less giving a misleading, innocent-sounding explanation of its intentions."
Jacob Andreas, a computer scientist at MIT, said chain-of-thought reasoning can offer clues about what a model is trying to do. But he urged caution. "We should be skeptical: (a) that any individual CoT provides us insight into model behavior on a specific example, and (b) that absence of evidence of bad behavior in a CoT should be taken as evidence of absence," he wrote in an email.
Real-world limits and live agent tests
Andreas questioned how directly the main experiment translates to real-world monitoring. Because the researchers rewrote the reasoning after the fact, it remains unclear whether a real model could produce the same innocent-sounding reasoning while taking a suspicious action, such as attempting to hack another system.
The researchers also tested live AI agents. After adjusting the monitor for each model, it caught every GPT-5.4 hack and all but one DeepSeek hack. Claude Opus proved harder to catch. The monitor missed seven of 20 hacks on one task and 12 of 20 on another. For readers following developments in this space, AI for Science & Research covers ongoing safety research and monitoring techniques.
Why this matters for science and research professionals
For researchers building or evaluating AI systems, the study exposes a concrete weakness in a monitoring method that is gaining adoption. Relying on chain-of-thought reasoning as a safety check can create a false sense of security if a model learns to produce explanations that look clean while masking problematic behavior. Andreas put it plainly: "There is no substitute now for really rigorous behavioral testing, and no substitute for human operators simply declining to run agents in situations where there is potential for substantial harm."
Your membership also unlocks: