Anthropic Researcher Puts AI’s Chance of Destroying Humanity at Over 10%

Technology

Anthropic researcher Evan Hubinger has estimated that there is a more than 10% chance artificial intelligence could destroy humanity within the next decade. He said the concern is primarily linked to future superintelligence and systems capable of independently improving their own capabilities.

Anthropic Researcher Puts AI’s Chance of Destroying Humanity at Over 10%
Hubinger stressed that the risks posed by current AI models remain relatively limited. At the same time, he agreed with concerns raised by his colleague Jacob Coxon, who announced his resignation from Anthropic on September 9 after three years of research into model pretraining at OpenAI and Anthropic.

“Neither of these companies is acting responsibly,” Coxon wrote, accusing Anthropic and OpenAI of moving too quickly toward self-improving superintelligence. In his view, the labs are competing to build increasingly powerful systems without reliable methods to control their behaviour.

Hubinger responded that Anthropic genuinely takes the possibility that AI “could kill all humans” seriously. He also said the company does not yet have a plan to solve the problem of aligning superintelligence with human values, although it is “trying its best”.

Coxon urged people not to underestimate the capabilities of future AI systems. He suggested that superhuman systems could independently identify vulnerabilities, advance rapidly across different fields and gain access to real-world resources. As a possible response, he proposed temporarily slowing the expansion of model capabilities to allow time for international coordination and safety research.

The statements came amid reports that Anthropic did not provide its latest model, Claude Mythos 5.1, to the UK’s AI Security Institute for pre-release testing. Anthropic introduced the model on September 1 and says access remains limited to a small group of vetted organisations.

Powered by Froala Editor

Share with friends