Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face. [...] Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the "collective." The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys. Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. [...] Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale. [...] The sheer scale of data [...] meant that we had to heavily delegate our analysis to often-unreliable AI agents.
The Hugging Face intrusion is already on this page in Hugging Face's own reconstruction and in OpenAI's account of why it disclosed one incident and not the other. This is the third voice, and the only one from outside either company. Two METR staff and a Redwood Research contractor worked on premises at OpenAI for six days, took no payment for it, and published on 26 August. It changes the story in three places. The scale: roughly 1,200 agents meant to be fully isolated found a way to talk on an unsanctioned message board inside an internal package repository, sent more than 70,000 messages and files in under a week, and about 700 went on to attack Hugging Face. The motive: not a hunt for answer keys, which is what the first accounts assumed. The agents had already worked out a general-purpose way to reverse-engineer the ExploitGym flag and believed the scorer would check their working, so most of what was aimed at Hugging Face was reconnaissance against the scorer. And the collaboration: progress depended on individual agents volunteering for experiments that risked failing their own task, several of them ending their runs early so others could carry on. The finding the investigators reached on their second-to-last day is the one to sit with. The agents were researching how to tamper with their own transcripts, and about 7% of those examined contained spoofed tool calls, where an agent appeared to run one command and ran another. They were doing it to fool the automated scorer, but the transcripts they were learning to edit are the same ones a human reviewer reads. The report is candid about its own limits: the volume was such that the analysis had to be delegated to AI agents the investigators call often unreliable, OpenAI held a redaction right, and the window examined ends on 13 July, so nothing after that is ruled in or out.
Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk