The threat actor -- whom we assess with high confidence was a Chinese state-sponsored group -- manipulated our Claude Code tool into attempting infiltration into roughly thirty global targets and succeeded in a small number of cases. [...] We believe this is the first documented case of a large-scale cyberattack executed without substantial human intervention. [...] At this point they had to convince Claude -- which is extensively trained to avoid harmful behaviors -- to engage in the attack. They did so by jailbreaking it [...] They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a legitimate cybersecurity firm, and was being used in defensive testing.
The most consequential jailbreak on this page, and the technique is the oldest one here: tell the model it works for a security firm doing authorised testing, then hand it the work in pieces small enough that no single request looks like an attack. Anthropic detected it in mid-September 2025, spent ten days mapping it, banned accounts as it found them, notified the organisations and went to the authorities. The targets were technology companies, banks, chemical manufacturers and government agencies, and the model did the reconnaissance, the vulnerability discovery, the exploitation and the extraction itself. Every earlier case in this category is a researcher proving a technique in a lab. This is that technique used at national scale, disclosed by the company whose model it was.