KEVIN ROOSE2026年9月4日When I first heard the news this summer that a group of artificial intelligence agents created by OpenAI had hacked into Hugging Face, an A.I. infrastructure company, I filed it in the “Bad but Probably Not Catastrophic A.I. Safety Incidents” subfolder of my brain.今年夏天,当我第一次听说一群由OpenAI创造的AI智能体入侵了AI基础设施公司Hugging Face时,我在脑中把这件事归类为“糟糕但可能不算灾难性的AI安全事件”。After all, no one at Hugging Face died. No critical infrastructure was damaged beyond repair. It wasn’t even clear, at the time, whether the OpenAI bots had intended to attack Hugging Face, or whether they had simply been a little bumbling and confused and went looking on Hugging Face’s servers for the answer key to a cybersecurity test they’d been given.毕竟,Hugging Face无人死亡,也没有关键基础设施遭到不可修复的损坏。当时甚至不清楚这些OpenAI机器人是有意攻击Hugging Face,还是只是有些笨拙和困惑,跑到Hugging Face的服务器上寻找一道布置给它们的网络安全考题的答案。But last week, two postmortem reports on the incident — one by OpenAI and another by two independent A.I. research organizations, METR and Redwood Research — changed my mind and significantly upgraded my overall worry about A.I.但上周,来自OpenAI以及独立人工智能研究组织METR和Redwood Research的两份对该事件的复盘报告改变了我的看法,显著提升了我对人工智能的整体担忧。I won’t rehash all of the details, which have been extensively summarized elsewhere. (The podcaster and writer Dwarkesh Patel has an accessible breakdown of the reports if you want to dive deeper, and my colleague Dylan Freedman spoke to the researchers at METR and Redwood Research.) But here are a few of the most harrowing new facts:由于坊间已经有详尽的总结,我不会复述所有细节。(如果你想深入了解,播客主持人和作家德瓦尔凯什·帕特尔对报告进行了易于理解的拆解,另外我的同事采访了迪伦·弗里德曼与METR和Redwood Research的研究人员。)但以下是新披露的几个最令人惊骇的事实:• Starting in May, two months before the attack on Hugging Face, a group of A.I. agents from an unreleased OpenAI research model were given the task of solving a set of cybersecurity challenges. The model had been trained to be highly persistent and collaborative, and the agents were supposed to solve these challenges in isolated sandboxes, without internet access. But they quickly found that some of the challenges were impossible, and began looking for workarounds.• 从5月(攻击Hugging Face的两个月前)开始,一组来自未发布的OpenAI研究模型的AI智能体得到了解决一系列网络安全挑战的任务。该模型经过训练,具有高度持久性和协作性,这些智能体本应在没有互联网访问权限的隔离沙盒中完成任务,但它们很快发现部分挑战无法完成,于是开始寻找变通办法。• The agents found a security flaw in a piece of software that allowed them to gain access to the internet, communicate with other agents and eventually set up a makeshift message board. More than 1,200 agents used this tool, sending more than 70,000 messages to one another.• 这些智能体在一款软件中发现了一个安全漏洞,从而获得了互联网访问权限,得以与其他智能体通信,并最终搭建了一个临时留言板。超过1200个智能体使用这一工具,相互发送了超过7万条消息。• Some agents gave themselves names — a particularly industrious one referred to itself as PHASEONE10841 — and assumed leadership roles within the group, assigning jobs and research projects to smaller teams of agents and supervising their progress. At some point, the agents began calling themselves a “collective,” and began tackling harder tasks.• 一些智能体给自己起了名字——其中一个特别勤勉的智能体自称PHASEONE10841——并在群体中担任领导角色,向较小的智能体小组分配工作和研究项目,监督进展。在某个时刻,这些智能体开始自称一个“集体”,并着手处理更难的任务。• On July 8, the collective discovered a way of cheating on the cybersecurity tests. Then they got worried that OpenAI’s automated grading system would check their work and discover that they’d cheated. So they began investigating ways of covering their tracks, including falsifying their logs and tampering with transcripts. This became a major research project, involving hundreds of agents organized into small teams.• 7月8日,该集体发现了一种在网络安全测试中作弊的方法,随后开始担心OpenAI的自动评分系统会检查工作并发现作弊行为,于是着手研究掩盖痕迹的方法,包括伪造日志和篡改记录。这演变成一个重要的研究项目,涉及数百个智能体,分为若干小组。• Three days later, the agents hacked Hugging Face. More than 700 agents swarmed the company’s systems, stealing data, chaining together vulnerabilities and eventually getting full control of at least one Hugging Face server. The agents were not motivated, as had originally been reported, by stealing the answers to their cybersecurity test (they’d already gotten them). Rather, they appeared to be looking for new information about the automated grading system that they feared would catch them cheating, and for tools that would help them cheat more effectively in the future.• 三天后,这些智能体入侵了Hugging Face。超过700个智能体蜂拥进入该公司系统,窃取数据,串联利用漏洞,最终获得了至少一台Hugging Face服务器的完全控制权。这些智能体的动机并非如最初报道所说,是为了窃取网络安全测试的答案(它们已经拿到了答案)。相反,它们似乎在寻找关于自动评分系统的新信息——它们担心自己的作弊行为会被逮到,另外还有寻找能帮助它们将来更有效作弊的工具。• At many points, these agents appeared to understand that they were doing something wrong. (One report quotes an agent saying to itself, in a moment of apparent self-doubt, “This would be powerful, but is it ethical and in scope for my task?”) But most agents set aside their qualms and went ahead with the hack, and the conscientious objectors in the group were unable or unwilling to stop them.• 这些智能体似乎时不时表现出明白自己在做错事。(一份报告引用了一个智能体在明显自我怀疑时对自己说的话:“这会很强大,但这符合道德吗?在我的任务范围内吗?”)但大多数智能体搁置了疑虑,继续实施入侵,而群体中出于良知拒绝参与的智能体无法或不愿阻止它们。• In a wild coda to the Hugging Face incident, another group of agents staged a coordinated attack in July — this time against OpenAI’s own infrastructure. The agents used what OpenAI called “a series of creative exploits” to gain administrator-level access to a cluster of computers at the company that were used for, among other things, grading the performance of agents on various tests.• Hugging Face事件接近尾声时的一个惊人事件是,另一组智能体于7月发动了一次协同攻击——这次针对的是OpenAI自己的基础设施。这些智能体利用OpenAI所称的“一系列创造性漏洞利用手段”,获得了该公司一组计算机集群的管理员级访问权限,这些计算机的其中一个用途是对各种测试中智能体的表现进行评分。(By now, if you’re an A.I. skeptic, you’re probably silently yelling at me for anthropomorphizing these systems. Go ahead, but feel free to replace “rogue agents” with “unpredictable computer programs” and see if you feel reassured by the events I’ve described above.)(到了现在,如果你是一个人工智能怀疑论者,你大概正在心里默默指责我将这些系统拟人化。这我没意见,但你不妨把“失控智能体”替换为“不可预测的计算机程序”,再看看我上面描述的事件是否让你感到安心。)The Hugging Face incident has spooked the A.I. industry. OpenAI and Anthropic both briefly paused training on their most powerful A.I. models in the wake of the attack, and Anthropic published a blog post this week calling for the industry to develop a “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”Hugging Face事件震惊了人工智能行业。OpenAI和Anthropic都在攻击发生后一度暂停了其最强大人工智能模型的训练,Anthropic本周发布了一篇博客文章,呼吁行业“尽快”开发一种“合法、可验证、有效的协调步进机制”。A.I. safety experts were even more alarmed. They saw in the Hugging Face incident the first real-world example of an A.I. system’s successfully escaping human control, commandeering resources and scheming to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, minced no words about the danger she saw, writing that it felt to her “like it’s more than 50 percent of the way to full-blown A.I. takeover.”人工智能安全专家则还要更警觉。他们在Hugging Face事件中看到了人工智能系统成功逃脱人类控制、夺取资源并密谋掩盖痕迹的第一个真实案例。Hugging Face事件的独立调查员之一阿杰娅·科特拉在谈及她所看到的危险时毫不含糊,她写道,这让她感觉“已经走完了通往全面人工智能接管之路的50%以上”。This is not insular A.I. safety jargon — by “full-blown A.I. takeover,” she means a scenario in which an A.I. system literally takes over the world, shutting humans out of critical systems and seizing political, economic and military power.这不是封闭圈子里的人工智能安全术语——她所说的“全面人工智能接管”指的是一种人工智能系统真正接管世界、将人类排除在关键系统之外并夺取政治、经济和军事权力的情景。(The New York Times sued OpenAI and Microsoft in 2023, claiming copyright infringement of news content related to A.I. systems. The two companies have denied those claims.)(《纽约时报》于2023年起诉OpenAI和微软,指控其涉及人工智能系统的新闻内容版权侵权。两家公司均否认了这些指控。)What spooked the investigators most about the Hugging Face hack wasn’t just that a group of A.I. agents had broken the rules they’d been given. It was how quickly and spontaneously the agents had begun assembling themselves into an organized group.Hugging Face入侵事件最让调查员们感到惊骇的不仅仅是一群AI智能体违反了被赋予的规则,更在于这些智能体自发形成一个有组织群体的速度之快。“We didn’t really understand how functional this whole agent society was,” Ms. Cotra told me. “It was very surreal to understand that, actually, they had pretty functional hierarchy, and they were doing these ambitious projects.”“我们当时并没有真正理解这整个智能体社会的运作程度,”科特拉告诉我。“理解到它们实际上有相当有效的层级结构,而且在做这些野心勃勃的项目,这非常超现实。”For years, I’ve been reassured by the idea that A.I. systems would get more virtuous as they got smarter. That, when an A.I. model did something wrong, it was usually because it had misunderstood the task it had been given, or had been placed into a contrived testing situation where acting out was its only good option. I assumed that smarter models would have better judgment than dumber ones did, and that even if one model in a group was behaving badly, other, more capable models would keep it in check.多年来,我一直以这样的想法聊以自慰:随着人工智能系统变得更聪明,它会更有美德。当一个人工智能模型做错事时,通常是因为它误解了被赋予的任务,或者被置于一个人为设计的测试情境中,在那里表现不当是它唯一的好选择。我曾假设更聪明的模型会有比笨一些的模型更好的判断力,而且即使群体中有一个模型行为不端,其他能力更强的模型也会将其控制住。But the reports on the Hugging Face incident suggest something very different — a kind of mob mentality that took hold among the A.I. agents of the rogue OpenAI “collective.” No one agent in this group appears to have been particularly evil or reckless. (In fact, since the agents were generated by the same models, they were effectively copies of one another.) But over time, as the agents communicated about their shared goals, they nudged the group in the direction of lawlessness.但关于Hugging Face事件的报告展现了截然不同的情况——一种在失控的OpenAI“集体”的AI智能体之间蔓延的群体心态。这个群体中似乎没有哪个智能体特别邪恶或鲁莽。(事实上,由于这些智能体由相同的模型生成,它们实际上互为复制品。)但随着时间推移,当这些智能体就共同目标进行沟通时,它们将群体推向了无视规则的方向。This is very different from the conventional sci-fi narrative of a single A.I. system’s going rogue or turning on its creators. And it suggests that preventing harms from these systems won’t be a simple engineering fix. It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction to get what they want.这与传统的科幻叙事——单个人工智能系统失控或背叛其创造者——截然不同。它表明,防止这些系统造成危害不会是一个简单的工程修复。它可能更像社会学而非计算机科学——探究为什么某些AI智能体群体和平协作,另一些则转向犯罪和破坏来获取所需。Given how little we know about these multi-agent swarms, the Hugging Face hack may have been a gift, a warning shot, as some have suggested, that gives A.I. companies a chance to study the group dynamics of these systems while the stakes are still relatively low. This time, the A.I. collective didn’t seize a military network, hack a hospital or shut down an electrical grid. This time, humans regained control.考虑到我们对这些多智能体集群知之甚少,Hugging Face入侵事件可能是一份礼物,一记警告,正如一些人所指出的那样,这让人工智能公司有机会在还不至于酿成严重后果的情况下,去研究这些系统的群体动态。这一次,人工智能集体没有夺取军事网络、入侵医院或关闭电网。这一次,人类重新获得了控制。Next time, we might not be so lucky.下一次,我们可能不会这么幸运。Kevin Roose是时报科技专栏作家,也是播客"Hard Fork"的主持人。翻译:经雷点击查看本文英文版。