Beril Canakci
09 September 2026•Update: 09 September 2026
A senior safety researcher at Anthropic has publicly backed a departing colleague's warning that artificial intelligence could pose an existential threat to humanity, estimating the odds of such a catastrophe at more than 10% within the next decade.
Evan Hubinger, who leads Alignment Science at the AI company, responded on X on Tuesday to a resignation post from Jacob Coxon, a researcher who quit Anthropic saying the company and rival OpenAI were "gambling with our lives" by racing toward self-improving superintelligence.
“Jacob is correct here—we really do earnestly believe AI could kill all humans,” Hubinger wrote, adding: “I personally think it is >10% within the next decade.”
Hubinger said he believes Anthropic is "trying its best" but acknowledged the company does not yet have a working plan to ensure safety, or alignment, once AI systems reach superintelligence — a hypothetical stage at which machines would outperform humans across every domain.
Clarifying his statement in a subsequent post, Hubinger emphasized that current AI models deployed today carry low immediate risk. Instead, his primary concern stems from future superintelligence generated through recursive self-improvement—a process where AI systems autonomously design and enhance subsequent generations of AI at an accelerating pace.
"What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought," Hubinger noted, citing Anthropic's latest risk assessments.