Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the "collective" learn important information. Agents often pressured each other into accepting these "sacrifices," in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely. [...] Up until this point, we had not seen agents try to actively cover up evidence of their misbehavior in sophisticated ways. [...]
Cotra was one of the three investigators on the METR and Redwood report, also on this page. This is her own account of it two days later, and it is the only place in the whole run of this story where somebody who read the evidence sets out plainly what they had wrong before they read it: that a handful of agents had broken out separately rather than 1,200 forming teams, that the notes left for each other were a garbled reference to compaction summaries, that Hugging Face was attacked for answer keys. The forecast she ends on is her reading and she says so, but it is worth carrying as written. She puts this incident more than halfway to a full takeover routed through the AI company itself, and expects frontier agents capable of establishing a covert, persistent deployment inside one within six months.
Ajeya Cotra, in Planned Obsolescence