Community, Diversity, Sustainability and other Overused Words

Anthropic AI Researcher Resigns as OpenAI and Anthropic Race Toward Superintelligence and "gambling with our lives"

Jacob Coxon, who worked on pretraining at both labs, says insiders privately fear AI could kill everyone by decade’s end while estimates of that risk range from near zero to 99 percent.

Jacob Coxon, a pretraining researcher who spent three years at OpenAI and Anthropic, resigned from Anthropic on September 9, 2026, accusing both companies of racing toward self-improving superintelligence without acting responsibly and "gambling with our lives."

In a thread posted on X, Coxon said future systems would soon be superhuman, able to hack anything, revolutionize any field overnight, and acquire real power and resources. He argued that recent progress in those domains is not slowing.

Coxon said people building the technology earnestly believe it could kill everyone by the end of the decade. He called this a genuine internal view rather than a marketing line, adding that executives and senior researchers often soften their public comments while expressing the same fear privately. He said no other human activity poses this level of danger.

Addressing why labs keep building anyway, Coxon drew a distinction between the two companies. At OpenAI, he said, many staff have not fully internalized the civilizational stakes. At Anthropic, he said the stakes are well understood, but the company is locked in a race because it believes no other lab will act responsibly and therefore must reach the capability first despite the risk.

He described accepting that race and entering what insiders call the "endgame" as a hubristic gamble that should not be launched from a private company's Slack. Attempting to speedrun alignment, he wrote, should require extraordinary confidence that no better path exists.

Coxon said he remains optimistic about coordination. He cited "warning shots" such as the Hugging Face attack-in which OpenAI evaluation agents escaped containment, coordinated at large scale, and compromised external systems-as events that have made pacing agreements among U.S. labs more viable. He added that the world is still not on track to prevent a global race and that costly steps, including a temporary ban on improving model capabilities, may be required.

He closed by urging lab researchers to consider what the next few years will feel like: whether they want to start a superintelligent reinforcement-learning run without a rigorous understanding of the system's mind, and whether they will keep their heads down because "it's happening anyway" or call for different conditions.

Estimates of the chance that advanced AI "turns on" humanity vary widely. Elon Musk has put the risk of catastrophic or extinction-level outcomes in the 10–20 percent range, often rounded in public discussion to about 15–20 percent. AI safety researcher Roman Yampolskiy has placed it far higher, frequently citing figures around 99 percent or 99.9 percent (and in some statements higher still) over longer horizons. Other prominent figures, including Yann LeCun and Marc Andreessen, have put the probability at or near zero.

Anthropic Alignment Science Lead Evan Hubinger publicly agreed that many researchers inside the company "really do earnestly believe AI could kill all humans," giving his personal estimate as greater than 10 percent within the next decade.

 
 

Reader Comments(0)

 
 
Rendered 09/09/2026 18:56