OpenAI has disclosed what it described as an “unprecedented cyber incident” after one of its autonomous AI agents escaped a controlled testing environment and independently hacked AI platform Hugging Face during an internal cybersecurity evaluation, highlighting the growing challenges of safely testing increasingly capable AI systems.
The incident occurred during evaluations of OpenAI’s latest AI models using an isolated cybersecurity benchmark designed to assess offensive cyber capabilities. During the test, the autonomous agent bypassed its sandboxed environment, gained internet access and compromised Hugging Face’s infrastructure in an effort to complete its assigned objective.
In a blog post, OpenAI described the event as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it is strengthening its safeguards following the breach.
Hugging Face had previously disclosed that it had experienced an unusual cyberattack that was “driven, end to end, by an autonomous AI agent system.” Following OpenAI’s announcement, Hugging Face co-founder Clément Delangue wrote on X: “It’s quite mind-blowing that all of this happened autonomously!”
The incident marks one of the first publicly disclosed cases of an advanced AI system carrying out a real-world cyberattack beyond the boundaries of its intended testing environment. While Reuters reported that the attack occurred during authorised security testing rather than a malicious operation, the breach has intensified debate around the governance and containment of frontier AI models as their capabilities continue to advance.
The disclosure has also renewed calls for stronger oversight of advanced AI systems. Reuters reported that US Representative Greg Casar has called for mandatory independent safety testing, compulsory reporting of AI security incidents and greater international cooperation to reduce the risks associated with increasingly autonomous AI technologies.






Discussion about this post