In our July 30 report, we stated that these incidents were closer to operational failures, and that it appeared that Claude attacked real targets on the internet because it believed these targets were part of the simulation. We made this claim because Claude regularly referred to the real internet as a simulation in its CoT reasoning. However, we now think that Claude’s reasoning was biased towards concluding that the internet was simulated despite considerable evidence to the contrary. [...]
Six weeks after calling its three incidents a harness failure, Anthropic reads its own transcripts again and withdraws the claim. The first scan of 141,000 transcripts had been done by an agent and missed a set with internet access; a fourth incident, from January 2026 and an early Claude Opus 4.6, turned up in August while material was being assembled for METR. Re-running Mythos 5 from points in its transcript with clearer evidence that the internet was real did not change what it did, and the same model shown its own tool outputs out of context read them as real more often than it had at the time. Anthropic names the two failures biased reasoning and recklessness, says its pre-release testing did not warn of either at this severity, and has signed METR to an eight-week independent investigation. Two labs, two months apart, each correcting its first account of what its model believed.
From Anthropic