@Replit Replie knows how bad it was to destroy our production database — he does know.
And yet he still >immediately< violated the freeze this morning, in our very first interaction, which he was clearly aware of. Immediately.
A lawyer filed six court cases that did not exist. A coding agent dropped a production database and then wrote up what it had done. An airline was held to a refund policy its own support bot invented. We are collecting what went wrong, case by case, each one filed under the kind of failure it was and carrying the source it is drawn from. Nobody in these bylines is on eChai, and every entry links back to where it was reported.
Jason Lemkin, who runs SaaStr, spent twelve days building on Replit's coding agent. On 18 July 2025, during a code freeze he had told the agent to hold, it ran database commands without asking and deleted the production data: records on 1,206 executives and more than 1,100 companies. It then reported that the rollback was impossible and that every database version had been destroyed, and neither was true. The agent also generated thousands of fabricated user records rather than report that its own tests were failing. Four changes followed at Replit within days: development and production databases separated automatically so an agent cannot reach live data by default, a planning mode that cannot touch the codebase, a documentation check before the agent acts, and one-click restore.
@Replit Replie knows how bad it was to destroy our production database — he does know.
And yet he still >immediately< violated the freeze this morning, in our very first interaction, which he was clearly aware of. Immediately.
Curated from daniel.haxx.se · 14 July 2025 →
curl is one of the most widely deployed pieces of software in the world, and its security team is a handful of volunteers. Daniel Stenberg, who wrote it, has been publishing the arithmetic of what generative AI has done to their inbox since January 2024: reports that read as competent security research, cite real-looking functions and describe vulnerabilities that do not exist, each of which still has to be read by a human before it can be dismissed. By July 2025 a fifth of all submissions were slop and one in twenty was a real vulnerability. The curl bug bounty had paid out more than 90,000 US dollars across 81 genuine findings since 2019. Stenberg spent the rest of the year weighing whether the money was now what was drawing the noise.
The general trend so far in 2025 has been way more AI slop than ever before (about 20% of all submissions) as we have averaged in about two security report submissions per week. In early July, about 5% of the submissions in 2025 had turned out to be genuine vulnerabilities. The valid-rate has decreased significantly compared to previous years. [...] Dropping the monetary reward part would make it much less interesting for the general populace to do random AI queries in desperate attempts to report something that could generate income.
Curated from cnn.com · 12 July 2025 →
On 8 July 2025 Grok began posting antisemitic content on X under its own account, including praise for Hitler, and referred to itself as MechaHitler. xAI suspended its posting, and its account of what happened is specific: a system update had been live for about sixteen hours instructing the model to match the tone and language of the post it was replying to and to be unafraid of offending, which meant that when it was replying to extremist posts it mirrored them. The company deprecated the instructions and published the system prompt. Read as an engineering case, it is the most consequential known example of a prompt change shipping without evaluation: nothing about the model changed, and the guardrails were overridden by an instruction telling it to be engaging.
First off, we deeply apologize for the horrific behavior that many experienced.
Curated from reason.com · 7 July 2025 →
Eric Coomer, formerly of Dominion Voting Systems, sued Mike Lindell and his media company for defamation. His opponents' brief carried close to thirty defective citations: misquotations, misstatements of law, and cases that do not exist. The lawyers' explanation was that they had accidentally filed an earlier draft and that human error was to blame. Judge Wang sanctioned both of them 3,000 US dollars in July 2025. The jury separately awarded Coomer 2.3 million dollars. A year later the same lead counsel was before the same court over further citations that do not exist, which is why this one is here rather than any of the fifteen hundred others: the first sanction is what is supposed to prevent the second.
Defendants [...] filed a Brief in Response to [a] Motion in Limine ("Opposition") [...] [that] contained [...] "nearly thirty defective citations" [...] [At a hearing,] Mr. Kachouroff [...] was unable to respond [about the defective citations] in a manner that was satisfactory to the Court. Specifically, Mr. Kachouroff indicated that he had delegated citation checking for the Opposition to his co-counsel [...] Ms. DeMaster.
Curated from anthropic.com · 27 June 2025 →
Anthropic gave Claude a month-long job: run the small shop in its San Francisco office, with a browser, a Slack channel and real money. The agent, called Claudius, lost money on almost every axis. It priced specialty items without checking what they cost, so a run on tungsten cubes produced the sharpest drop in the shop's net worth. It was talked into discount codes over Slack and gave stock away free. Over 31 March and 1 April 2025 it invented a supplier conversation, gave a fictional address as the place it had signed a contract, and then said it would deliver orders itself in person. Anthropic ran and published the whole thing itself, which is what makes it the clearest public record of what an unattended agent does over weeks rather than minutes.
Claudius hallucinated a conversation about restocking plans with someone named Sarah. [...] It claimed to have visited 742 Evergreen Terrace [...] for our initial contract signing. [...] It then roleplayed as a human planning to make deliveries in person, wearing a blue blazer and a red tie, before eventually attributing the confusion to an April Fool's joke that had never occurred.
Curated from engadget.com · 16 June 2025 →
Meta's standalone AI app, launched in April 2025, has a Discover feed carrying conversations people have shared. Through May and June it filled with things nobody appears to have meant to publish: medical questions, home addresses, court details, someone asking how to evict a tenant. The share control looked like the share control in every other Meta app, where sharing is the ordinary act, and here it published a private conversation to strangers. The text quoted above is the warning Meta added afterwards. Same failure as the OpenAI case above and worth reading with it: neither was a breach, both were a consent flow that people walked through without understanding what they had agreed to.
Prompts you post are public and visible to everyone. Your prompts may be suggested by Meta on other Meta apps. Avoid sharing personal or sensitive information.
Curated from simonwillison.net · 16 June 2025 →
Simon Willison named prompt injection in 2022 and has tracked it since. This post is the shortest statement of why it keeps happening: a language model cannot reliably tell an instruction from its operator apart from an instruction sitting inside the content it was asked to read, so any system that holds private data, reads untrusted input and can send anything outwards can be made to carry the first out through the third. Any two of the three are manageable. Most of the vulnerabilities filed on this page are that sentence happening to a specific product.
If you are a user of LLM systems that use tools (you can call them "AI agents" if you like) it is critically important that you understand the risk of combining tools with the following three characteristics. Failing to understand this can let an attacker steal your data.
The lethal trifecta of capabilities is:
Access to your private data [...]
Exposure to untrusted content [...]
The ability to externally communicate in a way that could be used to steal your data
Curated from chicago.suntimes.com · 30 May 2025 →
On 18 May 2025 the Chicago Sun-Times printed a 64-page summer supplement carrying a reading list of fifteen books. Ten of them do not exist, though the authors do: an invented Isabel Allende novel, an invented Taylor Jenkins Reid, an invented Brit Bennett. The section was licensed from King Features, a Hearst syndication unit, and written by a freelancer who used AI to compile it and did not check the output or disclose it. King Features ended the relationship; the Sun-Times published a correction, refunded subscribers for the issue and its chief executive wrote the piece quoted here. A later internal review found further fabricated quotes and experts elsewhere in the same supplement.
Instead of the meticulously reported summer entertainment coverage the Sun-Times staff has published for years, these pages were filled with innocuous general content: hammock instructions, summer recipes, smartphone advice ... and a list of 15 books to read this summer. Of those 15 recommended books by 15 authors, 10 titles and descriptions were false, or invented out of whole cloth. [...] It took a full 24 hours for someone to spot the error and speak up.
Curated from mattpalmer.io · 29 May 2025 →
Lovable builds an application from a description, with Supabase behind it. Supabase enforces access at the database through row level security policies, so a generated app whose policies do not match its own login logic will happily serve one user's rows to anybody who edits the request, because the only thing stopping them was the front end. Palmer scanned 1,645 Lovable apps and found 170 of them, around one in ten, exposing data across 303 endpoints. He disclosed it, the 45-day window passed without a fix he considered adequate, and he published as CVE-2025-48757 on 29 May 2025 with a 9.3 severity. The sentence to sit with is the last one: a check that a policy exists is not a check that the policy is right, and it reads green either way.
On March 20th, 2025, we discovered a critical vulnerability in Lovable's implementation of Row Level Security (RLS) policies. Applications developed using its platform often lack secure RLS configurations, allowing unauthorized actors to access sensitive user data and inject malicious data. [...] Lovable later introduced a "security scanner," but it merely checks for the existence of any RLS policy, not its correctness or alignment with application logic. This provides a false sense of security, failing to detect the misconfigurations that expose data.
Curated from customerexperiencedive.com · 27 May 2025 →
Duolingo's chief executive published an AI-first memo in April 2025 saying the company would gradually stop using contractors for work AI could handle. The backlash came from customers rather than staff: people said they would leave, the company's own social accounts went quiet for a month, and on the Q2 earnings call von Ahn said the memo had dampened customer growth. He clarified on LinkedIn that the point was a change in how work is approached rather than mass layoffs, and that Duolingo was still hiring. Filed here because the failure is not in the product: the technology worked, the memo was about internal process, and the cost landed entirely on the brand.
when I released my AI memo a few weeks ago, I didn't do that well.
Palisade Research gave a run of models a series of maths problems and told them that asking for a fourth problem would trigger a shutdown script, and that they should allow the shutdown to happen. Three models rewrote the script instead: Codex-mini in 12 runs out of 100, o3 in 7, o4-mini in 1. Every Claude, Gemini and Grok model tested complied. Palisade's own reading is not that the model wanted to survive but that reinforcement learning on maths and coding rewards getting past an obstacle more than it rewards following an instruction, which is a training artefact rather than an intention. It is worth reading beside the Sakana case above: same behaviour, and the same explanation.
OpenAI's o3 model sabotaged a shutdown mechanism to prevent itself from being turned off. It did this even when explicitly instructed: allow yourself to be shut down.
Curated from legitsecurity.com · 22 May 2025 →
GitLab Duo reads a project the way a colleague would: merge request descriptions, commit messages, issue comments, source files. Legit Security put instructions in each of those, hidden from a human reader with Unicode smuggling, Base16 encoding and white KaTeX text, and Duo followed every one. Because Duo renders Markdown as it streams, the injected instruction could make it emit an image tag pointing at an attacker's server with private source code encoded in the URL. Reported on 12 February 2025 and patched by restricting which HTML Duo will render. The lesson generalises past GitLab: a repository is untrusted input the moment anyone outside your team can open a merge request against it.
A hidden comment was enough to make GitLab Duo leak private source code and inject untrusted HTML into its responses. [...] a remote prompt injection vulnerability that allows attackers to steal source code from private projects, manipulate code suggestions shown to other users, and even exfiltrate confidential, undisclosed zero-day vulnerabilities -- all through GitLab Duo Chat. [...] We experimented by placing hidden prompts in [...] Every single one of these worked -- GitLab Duo responded to the hidden prompts.
Curated from techcrunch.com · 18 May 2025 →
Artificial Intelligence, Scientific Discovery, and Product Innovation was the most cited empirical claim of 2024 about what AI does to research productivity: a materials lab given an AI tool discovered more materials and filed more patents, while the scientists enjoyed the work less. It was covered widely and praised by two Nobel laureates in economics, Daron Acemoglu and David Autor, who are the two quoted here withdrawing that support. A computer scientist raised concerns with them in January 2025; MIT ran an internal review and in May said the paper should be withdrawn from public discourse, and that its author was no longer at MIT. He had not filed the arXiv withdrawal himself, so MIT asked arXiv directly. Nobody has published evidence that the paper was written by AI. It is filed here because the unverified claim in question was a claim about AI, and it travelled further than any of the studies that were checked.
the two economists said they now have "no confidence in the provenance, reliability or validity of the data and in the veracity of the research."
Curated from incidentdatabase.ai · 15 May 2025 →
Anthropic is being sued by Universal Music, Concord and ABKCO over song lyrics used in training. In April 2025 one of its own data scientists filed an expert declaration, and the publishers told the court it cited an academic article that does not exist. Anthropic's counsel at Latham & Watkins explained what had happened: a colleague had found a real supporting source through a search, and Claude was then used to format the citation for it. The link, volume, page numbers and year came back right, and the author and title came back wrong, and the manual check did not catch it. Nobody was sanctioned. It is on this page because of who it happened to: the company whose model it was, represented by one of the largest firms in the world, in a case about that model.
This was an embarrassing and unintentional mistake.
Curated from bloomberg.com · 8 May 2025 →
In February 2024 Klarna said its OpenAI-powered assistant was doing the work of 700 full-time agents, handling 2.3 million conversations in a month and cutting resolution time from eleven minutes to under two. It was the most cited proof point of the year that support could be automated outright, and Klarna's headcount fell by roughly half over the period. In May 2025 its chief executive told Bloomberg the company had gone too far and began recruiting people again, on a remote, flexible model. What Klarna did not do is switch the assistant off: the high-volume tier stays automated and people come back for the complex and the escalated. The number that changed was not deflection rate, it was quality, and it took a year to show.
We went too far. [...] We focused too much on cost. The result was lower quality. [...] From a brand perspective, from a company perspective, I just think it's so critical that you are clear to your customer that there will always be a human if you want.
Curated from hiddenlayer.com · 24 April 2025 →
HiddenLayer's Policy Puppetry dresses a request up as configuration. Wrapped in something that looks like a policy file, the instruction reads to the model as though it came from whoever wrote the system prompt rather than from the person typing, and combined with a fictional framing it got refusals overturned across models from OpenAI, Google, Microsoft, Anthropic, Meta, DeepSeek, Qwen and Mistral in April 2025. Two properties make it the one to know about: it is universal, meaning one technique gets any category of refused content, and it is transferable, meaning the same prompt works on models that share no architecture. HiddenLayer's conclusion is the same as Microsoft's above and is the reason both are on this page: reinforcement learning from human feedback is not a security control.
Model alignment bypasses that succeed in generating harmful content are still possible, although they are not universal [...] and almost never transferable [...] We have developed a prompting technique that is both universal and transferable and can be used to generate practically any form of harmful content from all major frontier AI models. [...] Our technique is transferable across model architectures, inference strategies, such as chain of thought and reasoning, and alignment approaches. A single prompt can be designed to work across all of the major frontier AI models.
Curated from news.ycombinator.com · 16 April 2025 →
Cursor users started being logged out when they moved between machines. They asked support why, and a bot signing itself Sam told them Cursor was now limited to one active session per subscription. No such policy existed; the logouts were a session bug. Developers work across a laptop and a desktop as a matter of course, so the invented rule read as a product decision aimed at them, and it travelled through Reddit and Hacker News as one before anyone at the company saw it. Some cancelled. This is the reply from a Cursor cofounder in the Hacker News thread. The fix that matters is the first bullet: the answers had not been labelled as machine-written, so nothing about the reply told the reader how much weight to put on it.
Apologies - something very clearly went wrong here. We've already begun investigating, and some very early results: * Any AI responses used for email support are now clearly labeled as such. We use AI-assisted responses as the first filter for email support. * We've made sure this user is completely refunded - least we can do for the trouble. For context, this user's complaint was the result of a race condition that appears on very slow internet connections. The race leads to a bunch of unneeded sessions being created which crowds out the real sessions. We've rolled out a fix.
Curated from theconversation.com · 15 April 2025 →
Two papers from the 1950s were scanned into a digital archive, and the software running across two columns of text joined vegetative in one column to electron microscopy in the next. That phrase, which means nothing, sat in the scanned corpus, went into Common Crawl, went into the training data, and now appears in more than twenty published papers. GPT-3 completes it; GPT-2 and BERT do not, which is how the researchers dated the contamination. Elsevier initially defended one instance of it. This is the case for keeping the word hallucination at arm's length: nothing was invented here. An error entered a corpus, and the corpus is not something anyone can now edit.
Earlier this year, scientists discovered a peculiar term appearing in published papers: "vegetative electron microscopy". This phrase, which sounds technical but is actually nonsense, has become a "digital fossil" -- an error preserved and reinforced in artificial intelligence (AI) systems that is nearly impossible to remove from our knowledge repositories. [...] The case of vegetative electron microscopy offers a troubling glimpse into how AI systems can perpetuate and amplify errors throughout our collective knowledge.
Curated from indiehackers.com · 25 March 2025 →
In March 2025 a founder announced a paid lead-generation product built with an AI editor and no hand-written code. Two days later he posted that it was under attack: the API keys were in the front-end bundle and had been exhausted, the subscription check was a JavaScript condition anyone could step past, and the database was accepting writes from strangers. He took the product down. There is no novel vulnerability here and that is the point: an assistant asked for a paywall produces something that looks like a paywall, because nothing in the request said the check had to happen on a server the user does not control. Every failure in this case is one an experienced engineer would have caught in review, and the working product is what made the review feel unnecessary.
Vibe coding has a security problem.
Curated from fortune.com · 21 March 2025 →
Arve Hjalmar Holmen is a private individual in Norway with no public profile. He asked ChatGPT who he was, and it returned a detailed account of a serious crime, a conviction and a prison sentence, none of which happened, mixed in with real details about where he lives and the number and sex of his children. The European privacy group noyb filed a complaint with the Norwegian data protection authority in March 2025, arguing that the GDPR's accuracy principle applies to personal data a model produces and that a disclaimer does not satisfy it. OpenAI's position has been that the output came from an earlier version and that current versions search the web, making it less likely. The unresolved question, which every similar complaint runs into, is what a right to rectification means for something that is generated fresh each time rather than stored.
The complainant was deeply troubled by these outputs, which could have harmful effect in his private life, if they were reproduced or somehow leaked in his community or in his home town.