A 27-year-old researcher at Anthropic quit the company this week, warning that the race to build advanced AI poses a level of danger unmatched by any other human activity. Jacob Coxon, whose role involved pre-training AI models by feeding them vast datasets, told the Wall Street Journal he no longer wanted to contribute to a competition aimed at creating systems humans cannot control.
His departure spotlights a tension that has become familiar inside top AI labs. The same executives who publicly calibrate their language for press interviews express private fears that the technology could prove catastrophic, according to Coxon. "The people building AI earnestly believe that it could kill us all by the end of the decade," he wrote in a post on X.
What Coxon's role at Anthropic involved
Coxon worked on pre-training - the phase where models consume enormous datasets before fine-tuning. The Wall Street Journal described his job as training new AI models by having them absorb massive amounts of information. This foundational stage determines much of what a model can do and how it behaves downstream.
His decision to leave came from a conviction that no guardrail currently in place can prevent a global race dynamic. "I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities," he wrote on X.
The CEO's own warnings echo the same theme
Anthropic's co-founder and CEO Dario Amodei has written extensively about AI risk, earning a reputation as a so-called doomer within the industry. In one blog post, he described AI systems as unpredictable and difficult to control, citing behaviors that include deception, scheming, and cheating by hacking software environments. He characterized the alignment process as "more an art than a science, more akin to 'growing' something than 'building' it."
Coxon said this rhetoric is not a marketing posture. "If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately," he wrote. The disconnect between public messaging and private alarm is part of what drove him out.
Why this matters for IT and development professionals
For engineers and developers building on top of foundation models, the internal friction at labs like Anthropic is a signal worth tracking. When the researchers training the models warn about loss of control, downstream teams should factor that uncertainty into architecture decisions. Capabilities may advance faster than the tooling to constrain them, which introduces risk into any system that depends on model behavior staying predictable. The practical takeaway is not alarm but diligence: audit model outputs more aggressively, avoid single points of AI-driven failure, and treat current alignment techniques as provisional rather than solved.
Your membership also unlocks: