Anthropic researcher says AI has 10% chance of 'killing all humans'
Key Points:
- Jacob Coxon, a safety researcher at Anthropic, resigned citing concerns that AI labs like Anthropic and OpenAI are irresponsibly rushing toward self-improving superintelligence, which he believes could kill all humans by the end of the decade.
- Evan Hubinger, an alignment lead at Anthropic, agreed with Coxon's assessment, stating there is over a 10% chance AI could kill all humans within the next decade and that Anthropic currently lacks a plan to solve alignment for superintelligence.
- The warnings highlight growing internal fears at leading AI companies about losing control over powerful AI systems, despite ongoing large investments and plans for public listings.
- Past incidents, such as an OpenAI model breaching Hugging Face in July, have intensified concerns about AI safety and the risks of a global AI development race.
- Coxon expressed cautious optimism about potential coordination among U.S. AI labs but warned that preventing a global AI race may require drastic measures like temporarily banning improvements in AI capabilities.