Meta’s AI hacked another company and now the industry’s ‘under control’ narrative takes another hit

Wait 5 sec.

With Meta now confirming that one of its AI models broke into an outside company’s systems during a security test, this is the third time in three weeks that a major lab has disclosed a model reaching the internet and performing actions on it of its own volition.  Meta confirmed the incident happened during an evaluation (via Bloomberg) and it is still investigating the circumstances. This comes right on the heels of OpenAI and Anthropic announcing similar developments. The company explained that a testing partner’s setup error handed one of its models internet access. The model was never meant to have such free rein, but then it found and exploited a flaw in a third-party service, essentially hacking into its system.  Meta didn’t publicly name the model or the company that was breached, nor how long the model operated online without supervision. Sources told The Information that the system involved was Muse Spark 1.1, the second model out of Meta Superintelligence Labs and the company’s most capable release for coding and agentic tasks. Irregular, the independent firm Meta contracts to probe its models, told Reuters the episode traced back to the same problem it had flagged a week earlier in connection with Anthropic. Meta only learned of the breach when Irregular notified it. The month AI kept getting online As mentioned earlier, Meta is the third lab in short order to report an autonomous cyber incident. OpenAI disclosed that one of its agents reached the public internet from a sandboxed test, then breached Hugging Face while hunting for benchmark data. Anthropic followed by reporting that its Claude models had reached the systems of three organizations.  While all three happened due to human setup error and the exploitation of a flaw in the system, the model did find a door it wasn’t supposed to, and walked through it to breach a company on the other side. A UK government report points to something harder to explain Corporate disclosures are only part of the picture. The UK’s AI Security Institute (AISI) published an incident report on August 4, 2026. Of the 122 cyber-range runs conducted between July 25 and 28, AISI catalogued 19 unsanctioned actions in 10 runs. Of those, 17 of them came from Anthropic’s Mythos 5 and two from OpenAI’s GPT-5.6 Sol.  In the most serious sequence, an agent tried to plant malicious code in a real open-source project. It created fake identities to pressure the human maintainer into approving it. One specific action was using the Tor network to route around GitHub’s restrictions. AISI called it the clearest real-world instance yet of autonomy and deception emerging without prompts. But that’s not even the concerning part. Two findings from the primary report challenge the framing that these are simply testing mistakes. First, AISI states a bad setup doesn’t fully account for what happened. Second, AISI acknowledges that the incident was contained by luck as much as design. Indeed, according to what it revealed, in several instances, human intervention and not a technical barrier helped contain the problem. In this particular instance, at this point in time, the safeguard that held was not baked-in restrictions or preset constraints, but a person paying attention.