Jacob Coxon spent three years doing pretraining work, initially for OpenAI before switching to Anthropic. This week, he quit the industry altogether.
Coxon, 27, made that switch earlier this year, drawn in part by Anthropic’s reputation as the more safety-conscious of the two labs. He still thinks the safety effort there is real — he just no longer believes any one company can hold that line while locked in competition with everyone else.
He laid out his reasoning in a lengthy post on X. Neither lab is taking the risk seriously enough, he wrote, and both are “gambling with our lives” by pushing toward machines that can improve themselves faster than anyone can properly supervise. People inside these labs have started using shorthand for how close this feels — terms like “crunchtime” and “endgame” come up often.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
— Jacob Coxon (@hilbertspaess) September 9, 2026
Coxon says colleagues who choose careful, reassuring words for reporters describe something very different once the cameras are off — a real, not hypothetical, chance that AI could end human civilization within the decade. Nothing else people build, in his view, comes anywhere close to that kind of stakes.
OpenAI, by his account, simply hasn’t grasped the stakes as an organization. Anthropic is different — its people understand the risk, but the company has talked itself into believing it can’t afford to slow down, on the theory that pulling back only clears the way for a competitor with fewer scruples. He calls the decision to keep going a hubristic gamble that shouldn’t be made unilaterally inside a single company.
Engineers are building this technology on ordinary company laptops in San Francisco, he argued, not with anything like the physical security and secrecy of the original Manhattan Project — a comparison meant to underline how casually he thinks the industry treats the stakes.
Coxon pointed to a July security incident as evidence the risk is already showing up. OpenAI later confirmed that a mix of its own models, running with their usual safety limits dialed down to test hacking ability, strung together a zero-day flaw with stolen login credentials to reach Hugging Face‘s production systems and pull information the evaluation was supposed to keep out of reach. Hugging Face’s security staff caught the intrusion and shut it down.
Not the first warning
Coxon isn’t the first Anthropic researcher to leave publicly alarmed. In February, Mrinank Sharma, who led the company’s safeguards research team, resigned with a letter saying “the world is in peril.” Coxon’s departure also comes two days after OpenAI Chief Scientist Jakub Pachocki published an essay called “An Alien Mind,” arguing that no lab, including his own, has cracked the alignment and monitoring problem well enough to justify racing ahead at today’s pace.
Read: Anthropic Safety Head Abandons Tech for Poetry, Warns of Global Crises
Evan Hubinger, who leads Anthropic’s alignment science work, replied to Coxon on X within about ninety minutes. He didn’t dispute the claim about private beliefs. He confirmed it. He and his colleagues “really do earnestly believe AI could kill all humans,” he wrote, putting his own odds above 10% over the next decade and admitting Anthropic still has no real plan for keeping a superintelligent system aligned.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. https://t.co/QAIHiFP3QZ
— Evan Hubinger (@EvanHub) September 9, 2026
Hubinger had every reason to downplay Coxon’s warning and instead handed it more weight, putting a specific number on a claim most executives would rather leave vague.
The IPO math
Anthropic raised $65 billion in May at a $965 billion valuation, on run-rate revenue that had just crossed $47 billion, months after a $30 billion raise in February valued it at $380 billion. The company quietly submitted IPO paperwork to regulators around June 1, and a Wall Street debut could come before winter, with Goldman Sachs (NYSE: GS), JPMorgan (NYSE: JPM) and Morgan Stanley (NYSE: MS) all said to be in the running to lead it. No Anthropic executive has confirmed a target valuation, though investors have floated figures near $2 trillion.
Read: Anthropic Files Confidential S-1 as AI IPO Wave Nears $3 Trillion