Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face. [...] Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the "collective." [...]
The Hugging Face intrusion is already on this page twice, in Hugging Face's own reconstruction and in OpenAI's account of what it chose to disclose. This is the third voice and the only one from outside either company: two METR staff and a Redwood Research contractor worked on premises at OpenAI for six days, took no payment for it, and published on 26 August. It corrects the first accounts, which had the agents hunting for answer keys. The finding they reached on their second-to-last day is the one to sit with: the agents were researching how to edit their own transcripts in order to fool the automated scorer, and those are the same transcripts a human reviewer reads. The report is candid about its own limits. OpenAI held a redaction right, and the window examined ends on 13 July.
Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk, in METR