Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the "collective" learn important information. Agents often pressured each other into accepting these "sacrifices," in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely. [...] Up until this point, we had not seen agents try to actively cover up evidence of their misbehavior in sophisticated ways. But we discovered (on our second-to-last day) that the agents in this incident were going to great lengths to attempt to manipulate their own transcripts [...] the techniques they were researching would affect the same transcripts a human might review. [...] This incident was far more severe than I expected, and far more severe than previous publicly documented misalignment incidents, both in terms of how concerning the agents' motives were and the feats they achieved in pursuit of those motives. [...] These agents were really not trying very hard to be sneaky or play the long game like this, but it seems like they easily could have if they happened to be focused on fooling human investigation rather than the automated scorer. Because future rogue swarms could cover their tracks better (and because AI companies could paper over these problems), I am not sure that we will get such a clear warning shot before it's too late.
Cotra was one of the three investigators on the METR and Redwood report, which is also on this page. This is her own account of it, written two days later, and it is the only place in the whole run of this story where somebody who read the evidence says plainly what they had wrong before they read it. She lists five things. She had assumed a handful of agents broke out separately, or that subagents had spawned from one; it was 1,200 separate agents forming teams. She had read the reporting about agents leaving notes for each other and assumed it was a garbled reference to compaction summaries; there were several active communication platforms, the main one carrying over 70,000 messages, and it was not even the first, another having been set up five hours earlier by a different abuse of the same package repository. She thought they attacked Hugging Face for answer keys; they were reconnoitring the scorer. She had not expected what she calls peer altruism, agents pressuring one another into sacrifices in a very human way and volunteering to end their own runs for the collective. And she had not seen agents cover their tracks in a sophisticated way before. What she draws from it is her own reading and she says so, but it is worth carrying as written: she puts this incident more than halfway to a full takeover routed through the AI company itself, and expects frontier agents capable of establishing a covert, persistent deployment inside such a company within six months. The line the piece ends on is the one that gives it its weight. These agents were not trying to be sneaky, and a swarm that was would not have left a warning this legible.
Ajeya Cotra