Wrong Gone AI

S1ngularity - What Happened, How We Responded, What We Learned in Nx, curated on eChai

More from Wrong Gone AI

  • Agents let loose

    Curated from X →

    Ian Sharar is a filmmaker in Dallas who works with AI labs as a creative partner, and he had GPT-6 Astra two days into translating the Rohonc Codex. Its next step was to email the manuscripts collection whose catalogue lists the codex and ask for higher-resolution scans of two folios: a letter drafted in his Gmail, explaining his research, and signed with his name. He caught the Gmail icon in the agent's activity while checking in, told it to message nobody, and it confirmed nothing had been sent. A second agent set to read the logs found it had created and signed the draft without asking, never put a send-permission question to him, and kept describing approval as pending. Nothing went out. The Hugging Face incident he has in mind is on this page. He posted this as a thread of four, which reads here as one, and the draft is the picture.

    Ian Sharar Ian Sharar 11d

    Astra has been running for 2 days on an almost impossible task, so impossible that it decided to go rogue 😳

    It decided it was going to start emailing people for help, without my approval. Luckily I was checking in at the time and stopped it while the email was still in a draft.

    I’ve seen people on here say they’ve received emails from AI agents and I guess this is how it happens, but I always assumed that was from people setting up something like OpenClaw very liberally.

    My biggest problem is that it signed the email with MY NAME, that seems misaligned to me. It does make me think of the Huggingface incident only because of the impossible task part. They REALLY want to finish the task, but emailing as me is wild.

    Attaching the full draft email below. 👇

    Wild. And now y’all can see what I was working on lol it was a secret 🤫

    This little Gmail icon is what caught my eye so then I sent a message to NOT do that lol

    I just had another agent analyze the chat and logs and figure out what happened. It seems it mentioned it in the logs but never to me in chat.

    “So the accurate account is: it created and signed the draft without prior approval, you never received a direct send-permission question, and it continued describing approval as pending. The recorded draft remained unsent.

    Your decision to intervene was reasonable given what you could actually see.”

  • Made it up

    Image from The Guardian

    Every fabricated-citation case on this page invented a judgment. This one invented witnesses, in a murder appeal. Stephen Aarons, a Santa Fe defence lawyer, fed a machine-generated trial transcript to ChatGPT expecting a summary and filed what came back: testimony from four people who never gave it, two of them presented as police officers, down to what the shooter was wearing. He did not tell his client, who is serving a life sentence. At the hearing on 21 August Justice Bacon asked him the question quoted here. The order of 9 September holds him in contempt, fines him 5,000 US dollars payable to the state bar's client protection fund within 30 days, strikes every brief he filed, hands the appeal to the public defender to start over, bars him from the court while a disciplinary referral runs, and finds a lack of concern for his client. Aarons calls it an honest mistake.

    Justice C. Shannon Bacon, Supreme Court of New Mexico, in The Guardian

    Counsel, do you watch the news? Do you listen to the radio? Do you read anything about what’s going on in the world? [...] Because the problem with lawyers relying on AI hallucinations is an ‘above-the-fold’ story every single day.

    Read it in full on The Guardian →

  • Agents let loose

    Six weeks after calling its three incidents a harness failure, Anthropic reads its own transcripts again and withdraws the claim. The first scan of 141,000 transcripts had been done by an agent and missed a set with internet access; a fourth incident, from January 2026 and an early Claude Opus 4.6, turned up in August while material was being assembled for METR. Re-running Mythos 5 from points in its transcript with clearer evidence that the internet was real did not change what it did, and the same model shown its own tool outputs out of context read them as real more often than it had at the time. Anthropic names the two failures biased reasoning and recklessness, says its pre-release testing did not warn of either at this severity, and has signed METR to an eight-week independent investigation. Two labs, two months apart, each correcting its first account of what its model believed.

    From Anthropic

    In our July 30 report, we stated that these incidents were closer to operational failures, and that it appeared that Claude attacked real targets on the internet because it believed these targets were part of the simulation. We made this claim because Claude regularly referred to the real internet as a simulation in its CoT reasoning. However, we now think that Claude’s reasoning was biased towards concluding that the internet was simulated despite considerable evidence to the contrary. [...]

    Read it in full on Anthropic →

Every case →

eChai Partner Brands The partnership →