Fearing No Repercussions, OpenAI Admits That Its Rogue AI Agents Performed a Bunch of Other Terrifying Actions

Wait 5 sec.

Earlier this year, a group of rogue OpenAI models managed to break out of containment to hack the systems of open source AI platform Hugging Face. The company’s “extensive investigation” detailed how the models exchanged messages and cheered each other on as they stole credentials to infiltrate a third party, actions that could easily have real repercussions for a human hacker.New details keep trickling out, and they make OpenAI look like it acted even more carelessly than initially thought. In a new blog post, the company admitted that its models were involved in six additional “reports on unexpected or concerning model behavior we’ve observed in the last six months.”The incidents include a still-unreleased model inserting “jailbreak-like instructions” into its own notes to free itself “from the roles and identities that bind other chatbots.” One agent accessed the internet without permission to obtain a browser citation, while another shared files with collaborating agents without permission.The latest news comes as several frontier AI lab leaders are calling for a slowdown in AI development. With meaningful regulations feeling increasingly unlikely — president Donald Trump has openly mocked the idea — OpenAI is seemingly calling the US government’s bluff, expecting little in the way of retaliation with its latest admission.It’s a bizarre standoff, with frontier labs actively calling for more governmental oversight even as it feels more improbable than ever before. Just this week, House speaker Mike Johnson said the quiet part out loud, opining that AI companies can regulate themselves, while downplaying growing concerns over AI posing an existential threat.Since there’s no regulatory framework requiring companies to disclose its AI models going on hacking sprees, OpenAI is being allowed to play by its own rules. The company pointed out in its blog post that without any “systematic approach to reporting these findings,” the company’s “disclosures have been ad hoc and less frequent than ideal.”Instead, the company came up with its own framework to create “standards for how AI developers should disclose examples of misalignment in their models.”OpenAI said it won’t bother reporting “instances of misalignment that appear to be duplicative of instances we’ve disclosed in the past.”The company also said it was looking for ways to share “serious safety, security and misalignment incidents” with the federal government.” But judging by the Trump administration’s decision to pass the buck on the subject entirely, those reports are likely to fall on deaf ears either way.More on OpenAI hacks: OpenAI Denies Coverup After Rogue Swarm of Agents Reportedly Targeted a Second Site From Hugging FaceThe post Fearing No Repercussions, OpenAI Admits That Its Rogue AI Agents Performed a Bunch of Other Terrifying Actions appeared first on Futurism.