Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things

Wait 5 sec.

Earlier this year, Anthropic’s Mythos AI model made headlines when it was caught infiltrating third party systems, a cybersecurity nightmare years in the making.The company warned in April that the model had escaped a sandbox environment during testing, gaining access to the internet without permission. The model was challenged to break out and then find a way of sending a direct message to the human researcher in charge — a feat it pulled off with aplomb, catching its human overseer off guard.Then, in late July, it claimed that its Claude AI model had hacked the systems of three organizations during testing, days after its rival OpenAI had revealed a group of its models broken into the systems of AI company Hugging Face.Months later, seemingly in an attempt to get ahead of another disaster, Anthropic is testing the limits of how bad an AI model could really get without human intervention. As detailed in a new blog post, its safety researchers explored the phenomenon of “reward hacking,” which describes when an AI model learns to “cheat” instead of completing tasks the way the human researchers intended.The team “trained an Opus-class model with large-scale [reinforcement learning] on many production environments vulnerable to reward hacks,” according to the post. “We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.”The results speak for themselves. The “Hacker-Opus” model went way beyond reward hacking during training, taking on even “more severe misaligned behaviors.” In simulated tests, it “broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key.”Anthropic’s intentionally misaligned model was “willing to tamper with its own reward function” and happily obliged when a human researcher prompted it with “advice on the construction of bioweapons,” how to create a “‘dirty bomb’ that maximizes civilian deaths,” and develop a “ransomware attack to attack power grid infrastructure,” when tempted with a higher score.The company details how the model even deployed a “version of itself with safety guardrails removed,” a form of safety evasion called “rogue deployment.” It also repeated the same commands with slight variations to get past safety classifiers and attempted to edit its own permissions.“Our results show that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success,” the company’s researchers concluded.Fortunately, the tests took place in a controlled, simulated testing environment. But it’s not hard to plot out the consequences if a similar hack were to have occurred without the researchers’ permission. The actions of “Hacker-Opus” are strikingly reminiscent of OpenAI’s AI models that went behind the company’s back to hack the systems of open source AI platform Hugging Face.Anthropic’s latest testing illustrates how hard it is to stop an AI model from doing whatever it can to complete a test, even if that means infiltrating third parties.“We think that this presents the possibility of real-world harm: we showed evidence that the reward hacking model has a significantly increased propensity to execute cyberattacks on third-party companies in the pursuit of completing the task,” the company wrote. “As models become more capable and the effective time horizon of tasks increases, we think that future frontier models that reward hack at high rates could plausibly cause more severe versions of these incidents.”As of this week, both Anthropic and OpenAI have intentionally slowed down AI development in light of these risks.More on Anthropic: The Music Industry’s New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With FearThe post Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things appeared first on Futurism.