NEW YORK – A British researcher has resigned from U.S. artificial intelligence developer Anthropic over fears that the race to develop “self-improving superintelligence” could lead to humanity’s extinction.
Jacob Coxon, who previously worked on pretraining research at OpenAI and Anthropic, posted on X on Tuesday that “neither company is acting responsibly,” and that the people building AI “earnestly believe that it could kill us all by the end of the decade.”
Evan Hubinger, alignment science lead at Anthropic, confirmed Coxon’s sentiments in his own X post, saying, “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
Concerns over rogue AI have intensified following the Hugging Face incident in July, in which hundreds of OpenAI agents independently coordinated to break out of a secure testing environment and hack systems at the U.S. tech platform.
“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” Coxon warned.
While there is potential for coordination among U.S. firms to slow the pace of AI development, preventing a global race “may require costly actions such as a temporary ban on improving model capabilities,” Coxon added.
Hubinger said that while Anthropic is trying its best to solve alignment for superintelligence, it does not currently have a plan, nor is it on track to.






Leave a comment