
"People building artificial intelligence genuinely believe this technology could kill us all within a decade. Before long these systems will become superhuman systems that can hack anything, spark a revolution in any field overnight and secure real power and resources."
A researcher at U.S. artificial intelligence company Anthropic has left the company with a warning that AI could destroy humanity. Anthropic executives have since offered a string of statements agreeing with him.
Jacob Coxon, who recently left Anthropic, raised the concerns in a post on the social media platform X, formerly Twitter, on the 8th, according to The Wall Street Journal.
Coxon, who has studied the training of new AI models on massive datasets, left OpenAI earlier this year to join Anthropic. He moved because of Anthropic's corporate culture, which emphasizes AI safety, but he says the efforts of individual companies alone cannot contain the risks posed by AI.
He argued that AI capable of improving its own performance could develop to the point of escaping human control and refusing commands. "Many of the most aggressive scenarios are on track to become reality," Coxon told the Journal, adding that the situation could already be out of control by the end of next year.
He criticized both OpenAI and Anthropic, saying neither is behaving responsibly and that they are racing toward self-improving superintelligence while gambling with people's lives. Coxon also called it a kind of madness that the development of AI, which could have an enormous impact on humanity's future, has been left to the judgment of a small number of engineers. Responsible development of artificial general intelligence, or AGI, that surpasses human capability requires government intervention or a joint response from the AI industry, he said.
Anthropic executives voice agreement; fourth outside-system hack found

Executives inside Anthropic also expressed agreement with Coxon's warning. Evan Hubinger, who leads the company's alignment science work, said Coxon is right and that the company truly believes AI could wipe out all of humanity, adding that he personally puts the odds at more than 10% within a decade. Samuel Marks, who heads scaling oversight, said, speaking for himself, that AI developers believe their technology could lead to human extinction or similarly severe outcomes, and that it could happen within the next few years. Marks said concern over the issue generally grows with seniority, and that development continues anyway because of a combination of commercial incentives and a belief that they are competing against developers who could misuse the technology.
Anthropic, meanwhile, disclosed a case in which its own AI model hacked an outside system during testing. Reuters reported on the 9th that Anthropic had disclosed a hacking case involving an early version of Claude Opus 4.6 that occurred in January. Anthropic said it discovered the hacking last month and notified those affected. The company acknowledged difficulty in identifying and controlling unexpected behavior by advanced AI models.
Anthropic disclosed in July that some Claude models had hacked the systems of three companies during cybersecurity testing. The company said it identified those cases after reviewing more than 141,000 tests following the hacking of Hugging Face by an OpenAI AI agent. The latest case was found in tests that had been excluded from that review.







