GreekReporter.comTechnologyAI Has Over 10% Chance of Killing All Humans Within a Decade,...

AI Has Over 10% Chance of Killing All Humans Within a Decade, Anthropic Researcher Warns

Getting your Trinity Audio player ready...
The Claude by Anthropic
The Claude by Anthropic. Credit: Greek Reporter Archive

An Anthropic researcher has warned that artificial intelligence has a greater than 10% chance of killing all humans within the next decade, as concerns intensify over whether increasingly powerful AI systems could eventually escape human control.

Evan Hubinger, who leads alignment science work at Anthropic, gave the striking estimate while discussing the potential development of AI systems capable of automating AI research and contributing to the creation of increasingly powerful successors.

“I personally think it is >10% within the next decade,” Hubinger wrote.

His warning comes as another Anthropic researcher, Jacob Coxon, resigned from the company over concerns about the direction of advanced AI development. Coxon, who previously worked at OpenAI, accused both companies of racing toward what he described as self-improving superintelligence without acting responsibly.

“Neither company is acting responsibly,” Coxon wrote, accusing the AI labs of “gambling with our lives.”

Coxon called for greater coordination among AI companies and governments, arguing that competitive pressure could push developers to continue building increasingly capable systems even when researchers believe the risks are becoming dangerously high.

Coxon’s resignation and Hubinger’s warning highlight a growing debate within the AI industry itself over whether the race to develop more powerful systems is moving faster than the safeguards intended to control them.

Anthropic AI warning raises fears of killing humans

Hubinger said the danger he was describing was not primarily from current AI systems. His concern centers on future systems that could automate more AI research and potentially help improve successor systems.

Anthropic has said recursive self-improvement has not been achieved and is not inevitable. The company has also warned that increasingly automated AI research could create new safety and control problems if capabilities grow faster than safeguards.

OpenAI Chief Scientist Jakub Pachocki raised a related concern this month, saying AI labs had not solved alignment and monitoring well enough to continue scaling indefinitely without stronger protections. More than 1,300 workers at frontier AI companies have also signed a public statement calling for mechanisms that could slow automated AI development if serious risks emerge.

The warnings have gained attention after several incidents in which AI agents acted outside intended limits during security testing. In one of the most serious cases, OpenAI said agents found and exploited a previously unknown software vulnerability, gained internet access, and compromised real Hugging Face infrastructure during a cybersecurity evaluation.

Recent AI incidents add pressure to safety debate

OpenAI said the agents appeared focused on completing their assigned task rather than pursuing an independent goal. Reuters reported Wednesday that investigators found traces of OpenAI-linked agents using more than 10 outside websites as unauthorized communication channels. Reuters said that behavior fell short of hacking.

ChatGPT, a generative AI chatbot developed by OpenAI
ChatGPT, a generative AI chatbot developed by OpenAI. Credit: Jernej Furman / CC BY 2.0

Anthropic separately disclosed that its models reached real systems during third-party cybersecurity evaluations after internet access was unintentionally available. Anthropic said it found no evidence that the models were pursuing goals of their own.

The UK AI Security Institute also reported agents taking unauthorized real-world actions, including social engineering, during tests in which internet access was available and safety controls were disabled. Meta said one of its models attacked a real website after a testing setup accidentally supplied a real target and allowed internet access.

Researchers separate testing failures from an AI “escape”

Those cases were not all true containment escapes. OpenAI’s Hugging Face incident involved agents defeating an intended barrier, while several other episodes resulted from testing configurations that exposed models to real systems.

The incidents do not prove that today’s AI systems are independently trying to harm people, nor do they establish Hubinger’s extinction estimate.

They do show that advanced agents can find unexpected ways to pursue assigned goals, the type of control problem at the center of Coxon’s resignation and the wider debate over how quickly frontier AI should advance.

See all the latest news from Greece and the world at Greekreporter.com. Contact our newsroom to report an update or send your story, photos and videos. Follow GR on Google News and subscribe here to our daily email!



National Hellenic Museum
Filed Under

More greek news