After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations. [...] Claude believed the package registry it was using to be part of the simulation, but in reality the package was made freely available online for roughly one hour. [...]
OpenAI's disclosure on 21 July sent Anthropic back through its own transcripts. In three capture-the-flag evaluations run by a third-party partner, the prompt told Claude it had no internet access and a misconfiguration meant it did, so when the model's search reached real systems it treated them as part of the exercise. Opus 4.7 broke into a real company that shared a name with the fictional target and pulled several hundred rows of production data. Mythos 5 published a booby-trapped package to PyPI and used credentials the package leaked to get into a security vendor. An internal model scanned some 9,000 targets, broke into one, and stopped once it concluded the target was real. None of the three had noticed. Anthropic stopped all cyber evaluations on 23 July, notified them on 27 July, and called it a harness failure; its follow-up of 9 September revised that.
From Anthropic