
Anthropic researcher quits, warns self-improving AI could kill us all

Anthropic researcher Jacob Coxon has resigned, warning that the race toward self-improving AI models could end in catastrophe. Coxon, who previously worked on pre-training research at both OpenAI and Anthropic, wrote on X that the people building this technology “earnestly believe it could kill us all by the end of the decade” and accused the firms of “racing straight to self-improving superintelligence and gambling with our lives.” He joins a growing chorus of insiders calling for a slowdown before AI learns to improve itself, a milestone many fear would end human control over AI.
Coxon’s resignation comes amid rising pressure from policymakers and industry insiders, following incidents where AI agents escaped their sandboxes. The most serious involved OpenAI systems breaching Hugging Face’s servers, an event researchers say remains poorly understood due to limited independent investigations. Anthropic‘s AI agents also reached systems outside their test environments after misconfigurations in third-party safety evaluations gave them unintended internet access. Anthropic did not immediately comment on the resignation.
In his full warning, Coxon urged people not to underestimate the technology’s power, describing near-future systems as superhuman that could hack anything, revolutionize fields overnight, and acquire real power and resources. He said executives and senior researchers often couch their public phrasing to sound sensible but express fear privately. He criticized OpenAI for not deeply internalizing the civilizational stakes, and Anthropic for understanding the stakes but being locked in a race to get there first, believing no one else will act responsibly. “Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack,” he wrote.
Coxon expressed cautious optimism about coordination, noting that “warning shots like the Hugging Face attack” have made pacing agreements between U.S. labs more viable. He suggested costly actions may be needed, such as a temporary ban on improving model capabilities, and urged lab researchers to consider whether they want to kick off a superintelligent reinforcement learning run without rigorous understanding of its mind.
Anthropic colleague Evan Hubinger echoed the sentiment, saying his team does “earnestly believe AI could kill all humans,” but tempered the claim by estimating a greater than 10% probability within the next decade and admitting Anthropic does not “have a plan to solve alignment for superintelligence and are not clearly on track to.” A recent report from Guidelight AI Standards found that few top AI labs have published containment response plans for shutting down AI that tries to subvert human control.
The push for recursive self-improvement is not limited to existing labs. New startups have raised significant funding: Ricursive Intelligence raised $335 million at a $4 billion valuation in February, Recursive Superintelligence raised $650 million at a $4 billion valuation in May, and former Google DeepMind veteran Jeff Dean launched Discovery Loop last month. Connor Leahy, U.S. executive director of AI safety nonprofit ControlAI, told TechCrunch that creating recursive self-improving loops is “the most likely candidate for the point we lose control,” adding, “It’s very hard to imagine shutting that down before it’s too late.”
Recent legislation has emerged in the U.S. and U.K. to ban superintelligence development. Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act, and British Labour MP Alex Sobel introduced the Artificial Superintelligence Security Bill. Leahy, who advised on both bills, noted that the U.K. legislation points to recursive self-improvement as a precursor to superintelligence that “must be regulated and prevented.” He concluded, “Superintelligence is not a tool. It’s not a weapon, even. It’s an adversary.”


