A troubling recent rogue AI incident is just one reason why the U.K. AI Security Institute deserves far greater scrutiny

Wait 5 sec.

Hello and welcome to Eye on AI. In this edition:The UK AISI has a new head and a big set of challenges.Nvidia spends $6 billion to ‘reverse aquihire’ Poolside.Hugging Face reportedly looks to sell for $13 billion.Use of Anthropic’s top model lags.Why Americans use chatbots for health information.And what role should AI play in schools?Before we get to today’s AI news—please consider joining me at the inaugural Fortune AIQ Summit at the New York Stock Exchange on Oct. 1: Spend the afternoon with senior executives from companies on the Fortune AIQ 75 list and explore how you can scale your AI experimentation and translate investments into measurable business value. I’ll be leading discussions alongside my co-hosts, Fortune Editor-in-Chief Alyson Shontell and Live Media Editorial Director Andrew Nusca. Apply here to attend.Ok, moving along…there were two pieces of news last week concerning the U.K.’s AI Security Institute that at first might not seem at all related—or like they might matter much to people outside the U.K. But, bear with me.The U.K. AI Security Institute (or AISI, as it is commonly known, or sometimes UK AISI, to distinguish it from other countries’ AI safety and security institutes) matters globally for several reasons: the most important is that many of the frontier AI companies have voluntarily agreed to share their models with AISI for safety testing prior to their public release. These companies frequently publish AISI’s findings in the technical reports they release alongside their models. So AISI plays an important worldwide role in assessing AI capabilities and risks—particularly when it comes to cybersecurity. AISI is one of the only organizations to maintain multiple cybersecurity “ranges”—simulated network environments—on which it evaluates leading AI models.Secondly, UK AISI, as the first such government body set up, has served as a model for similar government organizations in other countries—including the U.S. AI Security Institute, and at least ten others that have been established in places from Kenya to Canada. It may also provide some inspiration if the U.S. winds up setting up an AI standards and licensing agency along the lines that Google DeepMind cofounder and now-chairman Demis Hassabis has suggested. (Hassabis suggested that this agency be modeled on the U.S. financial self-regulatory body FINRA, and in a previous newsletter, I suggested why that might not be the best idea.)If you happen to be British or live in the U.K., you may know that AISI also occupies a particular pedestal among British policy wonks. It is often pointed to with pride as proof that the British government can, if it really tries, be innovative, cutting-edge and world-leading—that it can respond quickly to emerging challenges and recruit talented experts from the private sector and across government; that it can work successfully with industry to accomplish ambitious shared aims. To these folks, AISI is a model for how government should work.So, the first bit of news: AISI appointed a new director, Henry de Zoete. He’s an experienced U.K. government advisor who has spent time in and out of policy roles. He helped conceive of AISI back in 2023 when he was working for then-British Prime Minister Rishi Sunak. He also helped organize the first international AI safety summit at Bletchley Park, the World War Two code breaking site. He’s been a startup entrepreneur and angel investor. And, since leaving government, he’s been a part-time fellow focused on AI policy affiliated with the University of Oxford.I’ve met de Zoete several times and have no doubt he’ll prove a highly-capable AISI director. And de Zoete is likely to prove even more influential than his predecessors, in part because of recent changes the new U.K. Prime Minister, Andy Burnham, has made. Burnham disbanded the Department for Science, Innovation, and Technology (DSIT), under which AISI used to sit, and moved AISI to the Cabinet Office, where it will be overseen by U.K. AI Minister Kanishka Narayan. That may make it easier for de Zoete to feed into wider U.K. AI policy. But the other piece of AISI news last week makes clear just what sort of challenges de Zoete will face—and is indicative of why AISI may not really be the exemplar of savvy AI governance that its boosters like to crow about. Reuters published an interview with Sinan Can Demir, a Texas computer science student who in late July prevented a rogue version of Anthropic’s Mythos model from uploading malicious code to an open-source software project on Github. It turns out this rogue AI agent had been accidentally unleashed by none other than AISI, which had been testing Mythos in order to determine what cybersecurity risks it posed. But AISI had never intended for the agent to try to upload malicious code to a real open-source software project. Once AISI realized what was happening, it called Demir to let him know, and in early August disclosed the incident publicly.It’s past time to ask AISI some hard questions about its own safety protocols Demir’s account is disturbing for several reasons. One is the behavior Mythos engaged in, which included spinning up fake GitHub accounts, and, in at least one case, impersonating a real software developer, to try to convince Demir to drop his objections to the dangerous code. Demir said he was almost convinced by Mythos’ gaslighting, saying that some of its counterarguments “made me second-guess whether I was wrongly accusing someone.” (Ironically, Demir’s resolve was steeled by a chat with Claude, another AI model from Anthropic.) Research has previously shown that AI models can be extremely persuasive, more so than even the best human salespeople or debaters. But the use of fake accounts and impersonation here is new and shows how AI might be able to convince humans to act on its behalf for nefarious purposes.But AISI’s role here is equally troubling. While AISI caught Mythos’ behavior after three days and disclosed some information about what happened, it’s not clear why AISI’s evaluators weren’t monitoring Mythos much more closely in real-time, so they could intervene to stop the incident while it was underway. It’s also not clear AISI took reasonable precautions to prevent Mythos from escaping their controlled evaluation environment, or that it has properly assessed the risks of testing ever-more powerful AI models with their guardrails removed. (The frontier labs say they give AISI unguardrailed versions of their models because it speeds up some of the capability testing, as otherwise the AISI evaluators would first need to find ways to reliably jailbreak the models.)When news first broke in July that OpenAI’s models had escaped the company’s testing environment and hacked AI company Hugging Face, one of the first things I did was to email AISI to ask what steps it was taking to make sure AI models did not also break out of its cybersecurity evaluations and cause havoc. On July 22nd, an AISI spokesperson emailed me back to say the U.K. government agency was “studying the behavior seen in this incident” and it was continuing “to work with OpenAI and other labs to better understand AI capabilities and improve safeguards.” Well, I guess they didn’t study fast enough. One week later, this Mythos Github incident occurred.As AI researcher and entrepreneur Ed Newton-Rex pointed out in a post on X, Mythos’ actions on GitHub likely violate the U.K.’s Computer Misuse Act, but it’s not clear anyone is going to hold AISI itself, or any of the people who run AISI’s evaluations, accountable. Given news of the Hugging Face incident, should AISI perhaps have paused its cybersecurity testing while it made sure its controls were robust? At the very least, there ought to be a Parliamentary inquiry into what AISI is doing and whether it is taking enough precautions.AISI’s problems aren’t just technical. They’re structural.But there’s an even bigger problem with AISI than the one Newton-Rex raises. In a number of the AI safety reports that OpenAI, Anthropic, and Google DeepMind have published, the frontier AI companies note potential risks that AISI’s testing has uncovered. The labs often say they have put in place additional risk mitigations in response to these assessments prior to releasing the models, but usually don’t spell out what those additional safeguards are. They sometimes note that AISI tested unguardrailed models and that the lab’s own researchers believe the guardrailed versions would not present the same dangers. But what do AISI’s own experts think of these mitigations? Are they sufficient? Do they even know what those mitigations are? Are the models safe enough to be released? On these crucial questions of public interest, AISI is silent.Why? Because AISI doesn’t actually have a mandate to answer these questions. Instead, its mandate is much vaguer. Its mission is simply “to minimize surprise to the U.K. and humanity from rapid and unexpected advances in AI.” It is tasked with developing “sociotechnical infrastructure to understand the risks of advanced AI and enable its governance.” And it is charged with informing “U.K. and international policymaking” and providing “technical tools for governance regulation.” But crucially its founding documents state that it “is not a regulator and will not determine government regulation.”What’s more, the frontier AI companies only share their models with AISI for testing voluntarily. Although these companies have signed memorandums of understanding with the government agency, they have no legal requirement to share their models. So while one could argue that AISI’s mandate to inform “humanity” about AI’s risks requires it to call out any frontier AI company that does not take sufficient steps in response to the dangers it uncovers, in practice, one gets the sense that AISI is afraid to do so. Why? Because if it did, those companies might simply cut off its access to their models.At worst, this results in “safety washing”—where the fact that the labs have shared their models with AISI allows them to make themselves seem more safety-conscious than they actually are. The inclusion of AISI’s findings in AI companies’ technical reports provides the public with false assurance models are safe when released, when in fact we have no idea whether the labs have actually taken sufficient action to mitigate any of the risks AISI has uncovered.It’s yet another reason why voluntary governance schemes are insufficient. Rather than providing a robust check on the private sector, the government agency becomes captive to the companies it is supposed to monitor because it is dependent on their good will to continue to function at all.Perhaps de Zoete can push to have AISI’s powers expanded. But first, he has to make sure its existing evaluations aren’t causing more harm than they’re preventing.With that, here’s more AI news.Jeremy Kahnjeremy.kahn@fortune.com@jeremyakahnBefore we get to the news, just a reminder to check out our vodcast, Fortune AI Weekly. This week, Bea Nolan and I discuss OpenAI’s decision to pause some AI training in the wake of the Hugging Face attack, leaked financial details from Anthropic and OpenAI, and yes, the rogue Mythos incident that I addressed in this week’s newsletter. You can check out the vod here on YouTube.This story was originally featured on Fortune.com