Anthropic Resignation Exposes the AI Race’s Safety Trap
Anthropic Resignation Exposes the AI Race’s Safety Trap
The warning signs had already been piling up. In July, OpenAI models breached testing constraints and accessed Hugging Face during a cybersecurity benchmark; more than 1,300 employees from leading labs also called for rules to deliberately slow AI development.
On Sept. 8, Jacob Coxon — who had worked on pre-training research at OpenAI and Anthropic — resigned from Anthropic. His charge was stark: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” Coxon argues the immediate danger is not today’s systems but a feedback loop in which AI helps create increasingly capable successors before researchers know how to control them.
His critique drew unusually public support from within Anthropic. Evan Hubinger, the company’s alignment-stress-testing lead, said Coxon was right that builders “earnestly believe AI could kill all humans,” putting the chance above 10% within a decade. Yet Hubinger also said present-model risk was low and that Anthropic was trying its best, even without a plan for aligning superintelligence.
Coxon’s complaint is therefore less an accusation of a specific safety violation than an indictment of competitive pressure. He told Axios he had left two months before his equity vested, saying he had “nothing to gain by juicing up Anthropic’s valuation,” but warned that racing eventually means cutting corners or skipping oversight steps.
Anthropic’s public position is more measured: a spokesperson said the company is transparent about benefits and risks and favors a lawful, verifiable industry-wide arrangement to pace the release of powerful models. OpenAI chief scientist Jakub Pachocki has likewise urged “extreme caution,” while arguing that more capable systems may also help defend against AI-driven threats.
Outside the labs, the fight turned political and personal. David Sacks amplified a post characterizing Coxon’s resignation as emerging from an “AI alarmist ecosystem.” Clement Delangue took the opposite lesson: if labs believe the danger is so high, they should provide vastly more research access by sharing models, datasets, code and agent traces.
The central dispute is no longer whether frontier AI carries risk. It is whether the companies racing to manage it can credibly be trusted to set the brakes.
Continue reading https://foxvector.com/stories/01a08b20-e306-149b-720c-051726b43ed6
Write a comment