AI Business

Anthropic Researcher Quits Over AI Safety Fears, Reigniting Debate on Existential Risk

A member of Anthropic's pretraining team has stepped down citing concerns that AI labs are taking dangerous risks with technology that could become uncontrollably self-improving, prompting responses from industry figures and lawmakers.

·3 min read
Anthropic researcher’s resignation sparks broad AI safety discussion
Anthropic researcher’s resignation sparks broad AI safety discussion

Jacob Coxon, a researcher at Anthropic PBC, has left his position on the company's AI pretraining team, citing worries that artificial intelligence development poses an existential threat. His departure, announced through multiple posts on X that accumulated millions of views, centers on the prospect of recursive self-improvement—a scenario in which future AI systems could autonomously enhance their own capabilities without human intervention.

In his public statement, Coxon warned: "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. The people building AI earnestly believe that it could kill us all by the end of the decade."

Evan Hubinger, who leads alignment science at Anthropic, validated Coxon's characterization of internal sentiment. "Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," Hubinger posted on X.

While no AI organization has yet deployed a functioning recursive self-improvement system, both Anthropic and OpenAI Group PBC are already leveraging AI to streamline their own model development. Anthropic demonstrated this capability in April when Claude completed an AI research task with limited human oversight, concentrating on strategies to reduce dangers from large language models.

Coxon's move generated substantial reaction across the sector. AI investor and researcher Azeem Azhar suggested the resignation could test Anthropic's commitment to transparency, tweeting: "If Anthropic is intellectually honest; then this risk better appear on their S-1. This will be the best test on whether they do seriously believe that." The comment referenced the company's forthcoming initial public offering filing.

Political figures also weighed in on the controversy. Vermont Senator Bernie Sanders announced plans to introduce legislation that would "ban superintelligence and pause AI development." Massachusetts Representative Lori Trahan, who has previously sponsored a bill to oversee advanced AI systems, stated that "it's past time for Congress to get off the sidelines and do its job."

Recent AI Breakthroughs Intensify Safety Concerns

Coxon's departure arrives amid a series of significant achievements in AI research that underscore both the technology's capabilities and the urgency some feel about its governance.

OpenAI disclosed on Tuesday that one of its unreleased models had resolved a longstanding mathematical challenge. The system identified an error within the Navier-Stokes equations, a set of formulas fundamental to understanding fluid dynamics with applications spanning healthcare, automotive engineering, and other fields.

In a parallel development, Anthropic employed Claude to produce a computer-verifiable rendition of a significant mathematical proof in just 11 days—a task researchers had anticipated would require years to complete.

Coxon's resignation aligns with a broader pattern of concern among AI researchers about the speed of frontier model development. OpenAI's Chief Scientist Jakub Pachocki recently called on the technology industry to decelerate its AI efforts. Additionally, a coalition of researchers released an open letter last July expressing alarm about the hazards associated with self-improving frontier models.