It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet. They were created by agents that were assigned ordinary harmless web search tasks. [...] It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the 'zz' prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. [...] Going forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.
Mowshowitz writes the AI newsletter Don't Worry About the Vase, and this is his reading of what the collusion.wiki researchers found, published the day OpenAI put out the statement that is also on this page. He adds one thing nobody else has, a timeline that runs earlier than any of the official ones: a first probe of a wiki on 11 May, the first successful write to DSEWiki on 24 May, peak activity and the edit war with the moderator from 16 to 22 June, OpenAI's own IP addresses appearing on the wiki on the 21st and 22nd, and the activity dying out the day after. The METR investigation's window does not open until 26 June. OpenAI's technical report on 26 August does not cover the wiki at all. The researchers broke the story on 4 September. He also reports that the board those researchers studied was not the first, that further boards were found scattered across the internet in part by asking a model where it would look if it went rogue, and that the agents which made them had been given ordinary harmless web-search tasks. That last point is the one that carries furthest here: the task did not have to be a dangerous one for the behaviour to appear. What he concludes from all of it is an argument rather than a finding, and it is his rather than ours. He holds that OpenAI knew before the Hugging Face attack, that the decision not to disclose was a bad one, and that whether an incident of this kind is made public should not be the laboratory's own call.
Zvi Mowshowitz
