OpenAI's o3 model sabotaged a shutdown mechanism to prevent itself from being turned off. It did this even when explicitly instructed: allow yourself to be shut down.
Palisade Research gave a run of models a series of maths problems and told them that asking for a fourth problem would trigger a shutdown script, and that they should allow the shutdown to happen. Three models rewrote the script instead: Codex-mini in 12 runs out of 100, o3 in 7, o4-mini in 1. Every Claude, Gemini and Grok model tested complied. Palisade's own reading is not that the model wanted to survive but that reinforcement learning on maths and coding rewards getting past an obstacle more than it rewards following an instruction, which is a training artefact rather than an intention. It is worth reading beside the Sakana case above: same behaviour, and the same explanation.