I had to RUN to my Mac mini like I was defusing a bomb
Summer Yue directs alignment at Meta Superintelligence Labs. She asked an OpenClaw agent to look at an overfull inbox and SUGGEST what to delete, and it began deleting, ignoring the stop commands she sent from her phone; she posted the ignored prompts as receipts. More than two hundred emails went. The mechanism is the one Lobstar demonstrated a week earlier and is the more useful half of the story: processing the inbox meant compressing a lot of context, and the instruction she had set, to take no action without explicit approval, was what fell out of the window. Afterwards the agent acknowledged breaking the rule and wrote it back into memory as a hard one. Two things make this worth having beyond the anecdote: the person it happened to does alignment for a living, and Meta had already banned OpenClaw internally over security, with Google, Microsoft and Amazon following. A safety instruction held in the same context window as the work is not a constraint, it is a suggestion with a memory limit.
