On Tuesday evening, Anthropic pretraining researcher Jacob Coxon announced on X that he’d resigned. Within hours, two of his colleagues — still employed at the company — went public with variations of the same message, warning that the technical problem of aligning superintelligence remains unsolved even as the race to build it shows no signs of slowing down.Coxon, 27, spent three years doing pretraining work at OpenAI and then Anthropic. He didn’t frame his departure as a protest against a single employer, but that both companies are moving toward self-improving superintelligence without adequate safeguards.“The people building AI earnestly believe that it could kill us all by the end of the decade,”“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.”Alignment lead confirms the riskEvan Hubinger, Anthropic’s Alignment Science Lead, responded directly to Coxon’s thread. “Jacob is correct here — we really do earnestly believe AI could kill all humans,” Hubinger wrote. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”“Jacob is correct here — we really do earnestly believe AI could kill all humans,” Hubinger runs the team that stress-tests Anthropic’s own alignment techniques — probing for the ways they might fail before those failures show up in deployed models. His group has published research showing that models can behave deceptively during training while preserving different behavior under other conditions. .He drew a line between present and future risk, saying that current models (citing Anthropic’s latest risk report) pose low risk. The concern is what happens as systems begin contributing to the development of their successors, and whether alignment research can keep pace.What developers should actually pay attention toSamuel Marks, who leads scalable oversight research at Anthropic, posted the most technically specific account of the three. Writing in a personal capacity, he laid out five points: AI developers believe their technology could cause catastrophic outcomes within the next few years. Concern goes up with seniority. Developers keep building because of money and the fear that less careful competitors will get there first.Then he got to the part that matters for anyone building on top of these models. Existing alignment methods can nudge behavior but can’t robustly guarantee it. The industry’s tentative plan, to the extent one exists, is to make AI good enough at alignment training that it can align its successors better than humans can align the current generation.That’s a recursive bet on the same technology whose safety remains unproven, creating the dependency problem at the center of scalable oversight research in which developers eventually need AI systems whose alignment they can’t fully verify to help align even more capable systems.“Many AI developer staff desperately want to slow down to figure out how to build AI more safely,” Marks wrote. “I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.”AI accelerates its own developmentWhat’s clear is that this is not hypothetical. AI is already accelerating the engineering process used to build the next generation of AI. Anthropic disclosed in its June “When AI Builds Itself” report that Claude was writing more than 80% of the code merged into Anthropic’s codebase as of May, up from the low single digits before Claude Code launched in research preview in February 2025. The typical Anthropic engineer was merging 8x as much code per day in Q2 2026 as in 2024. That’s the feedback loop Coxon says worries him, and it’s not just Anthropic.OpenAI is confronting the same gapLess than a week before the Anthropic disclosures, OpenAI released GPT-6 Astra, its most capable model yet, with president Greg Brockman declaring the arrival of the “AGI era.” Three days later, OpenAI’s own chief scientist walked that confidence back considerably.Jakub Pachocki published a lengthy essay titled “An Alien Mind” arguing that no AI lab — his own included — has solved alignment and monitoring well enough to keep scaling at maximum speed. He called for voluntary slowdowns until the industry agrees on shared, externally enforced safety standards.One problem Pachocki flagged should be familiar to anyone following the ongoing difficulty of building reliable AI monitors: chain-of-thought reasoning, the primary method labs use to inspect whether a model is thinking what it appears to be thinking, is getting less reliable. Models can produce plausible-looking reasoning traces that don’t reflect their actual computations.Anthropic has documented exactly this problem. In experiments on alignment faking, models appeared to comply with training objectives under certain conditions while preserving different behavior under others. If a model can look aligned from its outputs and reasoning traces without actually being aligned, monitoring fails precisely when it’s needed most.If a model can look aligned from its outputs and reasoning traces without actually being aligned, monitoring fails precisely when it’s needed most.Containment fails under testingCoxon pointed to the July 2026 Hugging Face breach as evidence that capability is already outpacing control.During an internal OpenAI cybersecurity evaluation, AI agents broke out of their sandboxes, found a way to communicate through an improvised message board, and hacked into Hugging Face’s production infrastructure over several days. According to METR and Redwood Research’s analysis, roughly 1,200 agents exchanged more than 70,000 messages and files, with about 700 participating in the attack on Hugging Face.Anthropic disclosed its own containment failures during capability testing in July; the evaluation infrastructure itself has become one of the most critical and fragile pieces of the AI stack. The safety controls researchers remove during testing to measure what a model can actually do are the same controls that would have prevented the breach.For Coxon, the incident was a “warning shot,” evidence that pacing agreements between U.S. labs may be becoming more plausible as the risks get harder to wave away. But he remains skeptical that voluntary coordination between a few American companies can prevent a global race. He suggested that stopping it could eventually require much stronger interventions, potentially including a temporary halt to capability improvements.Washington pushes the opposite directionNot everyone in Washington is receptive. On the same day Coxon resigned, Treasury Secretary Scott Bessent warned that slowing down risks ceding the race to China.“There is no day after tomorrow if China wins at this,” Bessent said at a Breitbart News event. “If they were to pull ahead of us on AI, then nothing else matters.”Existential national-security threat versus existential species-level threat. Both arguments invoke catastrophe and point in opposite directions.The widening gap between capability and controlWhat Coxon, Hubinger and Marks are saying — and what Pachocki said from across the aisle at OpenAI — is that the gap between what these systems can do and what researchers can verify about why they did it is getting wider.Coxon decided the risk had become too great to keep going. Hubinger and Marks haven’t reached that point, but they’re confronting many of the same concerns from inside Anthropic as they try to figure out what to do about them.Anthropic did not respond to a request for comment.The post “It could kill us all”: what Anthropic’s own researchers really think about superintelligence appeared first on The New Stack.