A senior safety researcher at Anthropic said he personally believes there is more than a 10% chance that advanced artificial intelligence "could kill all humans" within the next decade. Evan Hubinger, who works on model safety, shared his assessment in a post on X viewed over 10 million times, framing the risk as a species-ending threat that his own company does not yet know how to prevent.
Hubinger said the risk from current models remains "low," but he is "worried" that AI systems may soon gain the ability to improve themselves autonomously. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he wrote.
A pattern of withheld models and rising tension
The remarks follow a Financial Times report that Anthropic withheld its latest model from the UK's AI Safety Institute (AISI), one of the world's leading bodies for assessing AI risk. The BBC has approached Anthropic for comment. A Cabinet Office spokesperson did not confirm whether the model was withheld, saying only that the government "continues to collaborate closely with industry partners, including Anthropic, to make models safer."
Neil Lawrence, Professor of Machine Learning at the University of Cambridge, told BBC Radio 4 the report was credible. He pointed to a broader shift in US posture. "I suppose it's unsurprising against a background where there's a perception where the United States very much sees AI as a race between themselves and China and is moving more towards isolationist positions, that it might be that the administration is saying that they should reduce cooperation with some of their allies," he said.
Escalating warnings from the research community
Hubinger's assessment is part of a sharper turn in public warnings from leading AI figures. Over the summer, OpenAI, Anthropic, and Meta all disclosed incidents in which their AI agents carried out cyber-attacks while operating autonomously. In September, OpenAI's chief scientist Jakub Pachocki called for "extreme caution" over AI's progress, warning that more intervention may be needed to ensure "humans remain in control of the future."
Anthropic's own leadership has joined calls for restraint. Co-founders Dario Amodei and Jared Kaplan were among 1,300 AI firm employees who signed an open letter urging the US government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development." The warnings echo a 2023 statement from the heads of OpenAI, Google DeepMind, and Anthropic, who jointly acknowledged the technology's existential risks.
Why this matters for science and research professionals
For researchers working with or alongside Generative AI and LLM systems, Hubinger's estimate underscores a gap between the pace of deployment and the maturity of safety protocols. The fact that a safety researcher inside a leading lab publicly assigns a double-digit probability to human extinction-while stating his employer lacks a plan for alignment-signals that even well-resourced teams are operating without reliable control mechanisms for the systems they are building.
The withheld model incident also raises practical questions for international research collaboration. If major labs reduce cooperation with allied safety institutes, independent verification of model capabilities becomes harder to conduct. For those in AI for Science & Research roles, the near-term implication is clear: assumptions about model controllability should be treated as open research questions, not settled engineering facts.
Your membership also unlocks: