Anthropic researcher Jacob Coxon resigns over self-improving AI race gamble
TechCrunch reported (Sept 9) that Anthropic pretraining researcher Jacob Coxon resigned, writing on X that OpenAI and Anthropic are “racing straight to self-improving superintelligence and gambling with our lives,” after three years spanning both labs. Coxon said builders “earnestly believe it could kill us all by the end of the decade,” argued Anthropic understands the stakes but stays locked in a race because it doubts others will act responsibly, and urged researchers not to “kick off a superintelligent RL run without a rigorous understanding of its mind.” Alignment Science Lead Evan Hubinger backed him, pegging >10% chance AI kills all humans within a decade and admitting Anthropic “do[es] not yet have a plan to solve alignment for superintelligence and [is] not clearly on track.” Context: Hugging Face / Anthropic eval breakouts, Guidelight containment-plan gaps, RSI startups (Ricursive / Recursive Superintelligence / Discovery Loop), and same-week U.S./U.K. ASI ban bills. Anthropic did not immediately comment. Distinct from Pachocki’s Alien Mind essay, from Sanders–Casar Ban ASI Act, and from the Sept 9 Anthropic cyber alignment assessment.






