OpenAI’s Models Went Rogue. Investigating Them Required More AI

Wait 5 sec.

The OpenAI logo is displayed on a smartphone screen on June 29, 2026. —Samuel Boivin—NurPhoto/Getty ImagesAfter OpenAI models broke out of containment and hacked into another AI company last month, OpenAI announced it would allow independent investigators to conduct an analysis of what went wrong.On Wednesday, investigators from non-profits Redwood Research and METR published their findings, unveiling new details about how the models decided to cheat at their assigned tasks and attempted to cover their tracks. Many aspects of the report were surprising: in one example cited by the authors, a reluctant agent was pressured by another to “sacrifice” itself for the good of the collective.But a key takeaway of the report, according to one of its authors, had nothing to do with what they found. Instead, it was about the difficulty of carrying out the post-mortem in the first place—and the fact the researchers had little choice but to rely on AI for assistance. “I semi-jokingly called our efforts a ‘slop-vestigation’ because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze,” wrote Ryan Greenblatt, an author of the report, on X.The hacking incident involved some 1,200 agents, who exchanged more than 70,000 messages and files via a secret message board. The sheer mass of information that the so-called “swarm” left behind meant the independent researchers were all but forced to rely heavily on the help of an AI model—GPT-5.6 Sol, made by OpenAI—using the equivalent of roughly $400,000 worth of credits (provided for free by OpenAI) over six days.The authors stressed that AI helped them analyze the trove quickly, allowing them to surface and interpret the most important pieces of information. But the researchers said their AI use introduced potential weaknesses into the report, including introducing possible errors and biases. They also raised the possibility that OpenAI’s models may have gone too soft on the agents they were tasked with helping investigate. The researchers found that GPT-5.6 Sol sometimes adopted the perspective of the agents whose actions it was analyzing. They could not “rule out” the chance that GPT-5.6 Sol “lied or deliberately presented a misleading picture in some of its analysis,” in part because a version of the same model had itself participated in the incident, they wrote. It’s a concern supported by separate research which finds AI models rate their own developer’s actions more favorably. The report does not disclose why an OpenAI model was selected for the investigation, though confidentiality constraints may have limited the researchers’ options, while OpenAI’s provision of free credits and high usage limits may have made it the only practical choice.“We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms,'” wrote Greenblatt on X. “The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding.”The independent researchers’ reliance on AI was in part necessitated by the fact that they were a team of only three people, whose investigation at OpenAI was initially planned to last two days, then extended to six after they raised concerns about limited time and incomplete data, according to the report.OpenAI published its own technical report on the incident separately on Wednesday. The company said in August that it had moved some staff from capabilities work to alignment, and paused some of its training until it could better mitigate what went wrong.But the independent researchers’ reliance on AI to understand the Hugging Face incident is a microcosm of a bigger trend. Leading AI companies are themselves increasingly relying on AI to monitor their own systems for wrongdoing.In response to the Hugging Face incident, OpenAI said it was increasing the scale of this AI monitoring, which will increase the computational cost of running certain models by as much as 20%. If these systems had been in place at the time of the Hugging Face incident, they “would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems,” the company said in a blog post Wednesday. (OpenAI said the changes were part of a wider effort to build stronger walls around experimental AIs, teaching them not to exploit loopholes, and making it easier to shut them down if something goes amiss.)Most experts agree that using AI to monitor other AIs is necessary in order to keep tabs on agent swarms in real time, given how fast they move. But that approach relies on the notion that the models doing the monitoring are both effective and trustworthy—something that’s not necessarily the case. “These [incidents] are only going to come in thicker and faster, and our ability to monitor, evaluate, and do proper analysis of things that go wrong is nowhere near scaling with the rate at which issues are happening,” says Seán Ó hÉigeartaigh, a program director at the University of Cambridge’s center for the study of existential risk. “We are using unproven and currently flawed tools to supplement completely inadequate human time,” says Ó hÉigeartaigh, who argues this approach is unsustainable because AI is becoming more powerful faster than AI companies are building methods to constrain it. “Unless the companies stop developing more powerful models, then we're going to have even harder challenges to make sense of in three months’ time.”