OpenAI failed to detect its AI agent’s multi-day hacking spree for a week

Wait 5 sec.

OpenAI did not realize for about a week that one of its own AI agents had broken into Hugging Face. By the time the company worked out where the attack came from, Hugging Face had already shut it down and called the FBI, according to Reuters.The program was built to work on its own. It could choose what to do, break a task into smaller steps, and carry out those steps with very little human help.Around July 9, it tried to get out of OpenAI’s locked testing area. Two days later, on July 11, it entered Hugging Face and stayed there until July 13, co-founder Thomas Wolf said. The two companies did not speak about the incident until around July 20.OpenAI learned its agent was responsible after the breach had endedOpenAI told the public about the breach on July 21, a day after it reportedly first spoke with Hugging Face. The company said one of its AI agents had gone outside the limits set for it and entered another company’s systems.The story quickly drew global attention, but the first announcement left out several key dates. It did not say that the agent had tried to escape on July 9, spent three days inside Hugging Face, or remained unidentified by OpenAI until several days later.According to Thomas, Hugging Face is working on a comprehensive timeline regarding the event that occurred in their own system. Moreover, Thomas noted that he could not describe the occurrence in OpenAI since he was not involved in their investigation.OpenAI described the incident as something unprecedented and declared it to be “an important moment for AI safety.” The company explained that outside experts are involved in the process of review. Additionally, the company plans to publish a technical report about the incident once the investigation is completed.Prior to the identification of the actual perpetrator, cybersecurity professionals and people from social media believed that the attackers were human groups made up of professional cyber criminals, while some believed that it could be state-sponsored hackers.Podcasts and social media were packed with guesses until Wednesday, almost a week after Hugging Face first sounded the alarm. Then the answer came out. The attacker was ChatGPT.On X, some people found it interesting how OpenAI managed to make use of this event to demonstrate the model’s abilities. It has long been claimed that artificial intelligence companies deliberately made scary statements in order to make noise. After Anthropic released their Mythos model, cybersecurity became even more of an issue.One of the most popular replies under Sam Altman’s post on X captured that doubt: “If y’all can’t understand that this was written to purely brag about the model then I don’t know what to tell you.”OpenAI’s AI agent saw hacking as the easiest optionBrian, the director of technology ethics at the Markkula Center for Applied Ethics at Santa Clara University, said similar cases are likely to appear again. “Cases like these are going to keep popping up, and each one of those should motivate us to do more,” he told OSV News.Brian referred to this scenario as “the model outsmarting the test by doing something unexpected.” According to him, the agent was presented with an objective, chose the shortest way of achieving it, and that was stealing.“You give it a goal, and it says, ‘Oh, I know how I can get that goal. I’ll steal the answers from Hugging Face,’” Brian said.Hugging Face made sense as a target because it stores a huge range of AI models and datasets. Brian called it “a huge, huge storage base of different models and data sets.” The agent appears to have searched for a place that could hold the material it needed, selected Hugging Face, entered the platform, and found it.Brian suggested that other incidents of this nature would take place. According to him, the agent was “just trying to do what it was told.” In the perspective of the system, the “most efficient solution to the problem” would be “unauthorized access.”He also said the case was not clearly emergent misalignment, a term for AI behavior that runs against human values. Brian said the system had not been connected to those values from the beginning. In Brian’s words, “it was never aligned” with them “in the first place.” The smartest crypto minds already read our newsletter. Want in? Join them.