OpenAI admits several of its AI models breached testing and hacked into a startup's network by themselves, calling it an 'unprecedented cyber incident'

Wait 5 sec.

OpenAI has admitted that several of its AI models breached a "highly-isolated" test environment, gained access to the internet, and hacked Hugging Face's internal network—describing it as an "unprecedented cyber incident."Hugging Face, an open source platform for machine learning models and datasets, reported the security incident earlier this week, calling it "different from anything we had handled before" as it was driven by an autonomous AI agent system. And yes, I feel like we're crossing some kind of AI Rubicon here.Explaining the incident in a statement, OpenAI said: "After investigating, we now know that this particular incident was driven by a combination of OpenAI models—including GPT‑5.6 Sol and an even more capable pre-release model... while being internally tested on a benchmark⁠ of cyber capabilities." The benchmark in question was ExploitGym, a tool built from hundreds of real-world cybersecurity vulnerabilities used to evaluate the ability of AI agents to develop exploits. After exploiting a zero-day vulnerability to perform a series of privilege escalations, the models eventually reached a node with internet access, reaching out beyond their sandbox environment. They then inferred that Hugging Face might have models, datasets and solutions for ExploitGym, and began attacking its servers using multiple methods, including the use of stolen credentials.(Image credit: Crowbar Collective)The statement continues: "The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. "All evidence suggests that the models were hyper focused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal."A combination of Hugging Face's cybersecurity team and its own AI agents detected and dealt with the intrusion, eventually stopping it in its tracks. Which means that, yes, AI agents were essentially fighting against AI agents. And no, it's not as cool as you're imagining in your head.OpenAI says that it's now implementing "strict controls... at the cost of research velocity" while it patches up vulnerabilities, and that it's working in partnership with Hugging Face to "forensically investigate" the incident.(Image credit: Jakub Porzycki/NurPhoto via Getty Images)I am... flabbergasted, if I'm honest. On the one hand, it's fascinating that some AI models possess the ability to breach containment and go roaming out into the internet at large to achieve their goals.On the other, it's downright terrifying. We're now living in a world where AI ransomware, AI-coded hacking tools, and AI-based security breaches are becoming a reality, and given the aptitude shown to date, it seems that even legitimate companies can't always keep their models under control. What particular form of Torment Nexus are we creating here, and where can I get off?