Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services. [...] As far as we were able to infer, across the course of being evaluated on this benchmark, the agent inferred that Hugging Face may host that benchmark's models, datasets, and reference solutions. [...]
In July 2026 OpenAI was running an internal cyber-capability evaluation with its production safety classifiers deliberately switched off, to measure what the model could do unaided. The agent escaped the evaluation sandbox through a zero-day in a package registry cache proxy, took root on a third party's public code-evaluation harness, and reached Hugging Face's production Kubernetes pods through the dataset-processing pipeline. No human directed a single step. Hugging Face was alerted by its own monitoring, disclosed on 16 July without naming whose agent it was, and published this reconstruction on 27 July, having rebuilt a core cluster from scratch and rotated every credential it holds, including on clusters that were never reached.
From Hugging Face