An Anthropic researcher just resigned in public, and the industry's reaction says more than his warning did
Anthropic researcher Jacob Coxon resigned over AI safety concerns and went viral. Here is what he actually said, what Anthropic's own alignment lead said back, and what remains unverified.

Quick answer
Jacob Coxon, a 27 year old pretraining researcher who worked at both OpenAI and Anthropic over three years, resigned from Anthropic on September 9, 2026, and posted a seven part message on X warning that frontier AI labs are racing toward self improving superintelligence without adequate safeguards. He gave up unvested equity to speak publicly. Anthropic's own alignment lead, Evan Hubinger, responded that Coxon is correct that people inside Anthropic take the risk seriously. Coxon has also said he personally has not seen Anthropic compromise safety to beat competitors. Both things are true at once, and that tension is the actual story.
On the evening of September 9, Jacob Coxon posted on X, under his longstanding handle @hilbertspaess, that he had resigned from Anthropic. His post opened plainly: he had spent three years doing pretraining research, first at OpenAI, then at Anthropic, and he did not believe either company was acting responsibly.
The post spread fast, reportedly reaching tens of millions of views within a day, amplified by a Wall Street Journal interview published the same week. Coxon gave up equity that had not yet vested in order to speak on the record rather than staying quiet until his position was secure.
What Coxon actually said
Coxon's argument is not that Anthropic's safety work is fabricated. It is that competitive pressure between AI labs is making that safety work insufficient, regardless of how genuine it is.
In his interview with Axios, Coxon put it directly: being under pressure to race forces labs to cut corners or skip steps in oversight. He also offered a more complicated observation that got less attention than his headline warning, that some of the industry's caution is itself distorted by what he called excessive paranoia about OpenAI and about China, paranoia that can end up justifying pushing capability forward faster rather than slower.
Coxon's central technical concern is recursive self improvement, the idea that once AI systems can meaningfully improve their own training and architecture, progress could compound faster than human researchers can monitor or intervene in it. He has called for coordination among labs and a temporary pause on efforts specifically aimed at improving model capability, as distinct from other AI research.
What Anthropic's own people said back
This is the part worth being precise about, because it is easy to read as either a cover up or a confirmation, and it is neither cleanly.
Evan Hubinger, who leads alignment work at Anthropic, responded publicly rather than through a spokesperson statement, saying Coxon is correct that people inside the company earnestly believe advanced AI could pose a severe risk to humanity, and that he personally would put meaningful odds on a catastrophic outcome this decade. That is a striking thing for an alignment lead at a major AI lab to say on the record, and it happened in response to Coxon, not as a private admission that later leaked.
Coxon has also said, distinctly from his broader warning, that he personally has not seen Anthropic compromise its safety commitments to outlast competitors. That detail is often dropped from summaries of his resignation, but it matters: his argument is about systemic and competitive risk across the industry, not a specific allegation that Anthropic broke a promise internally.
Reporting following the resignation also noted that Anthropic is facing separate questions about its engagement with the UK's AI Security Institute, the government body that evaluates advanced AI systems. That is a live, developing thread and not yet a settled part of the Coxon story itself.
Why this is different from earlier high profile AI safety exits
Coxon's departure has been compared to earlier resignations at OpenAI, including Jan Leike's 2024 departure, where Leike directly alleged that safety work had lost out to product priorities internally. Coxon's framing is different. He is not alleging that either company broke a specific internal commitment. He is arguing that the structure of competition between labs, each individually acting in reasonably good faith on safety, can still produce a bad collective outcome. That is a more structural claim, and a harder one to fact check against any single company's internal conduct.
What remains genuinely unverified
Coxon's most viral line, that AI could kill everyone by the end of the decade, is a prediction, not a documented event. It reflects real, stated concern from people close to frontier model development, including Hubinger's own public odds. It is not a confirmed technical finding, and treating it as one in either direction, dismissing it as hype or repeating it as settled fact, would misrepresent what has actually been established here.
What is established: a credentialed researcher who worked inside two leading labs resigned publicly, gave up compensation to do so, and a senior safety researcher at the company he left corroborated the sincerity of his concern rather than disputing it. What remains open is everything about timelines, likelihood, and what specific policy response would actually address the risk Coxon describes.
FAQ
Who is Jacob Coxon?
A 27 year old AI researcher who spent three years doing pretraining research, first at OpenAI, then at Anthropic, before resigning from Anthropic on September 9, 2026.
Did Anthropic dispute his claims?
No. Anthropic's alignment lead, Evan Hubinger, responded publicly that Coxon's characterization of internal concern is accurate. Coxon himself has separately said he has not personally witnessed Anthropic compromise safety commitments.
Is this the same as previous AI safety resignations at OpenAI?
Not exactly. Earlier departures, such as Jan Leike's from OpenAI, alleged a specific internal shift away from safety priorities. Coxon's argument is broader and structural, that competitive pressure across the whole industry undermines safety work even where it is genuine.
What is Coxon actually asking for?
Coordination among AI labs and a temporary pause specifically on efforts to increase model capability, not a halt to AI research generally.
Sources
Ask MAMBO
Have a plain-English question about this topic? Send it in and we may answer it in a future guide.
Ask a question