NEWS
Anthropic Researcher Walks Out Over Superintelligence Race
Jacob Coxon left OpenAI for Anthropic, then quit in September. The safety lab’s alignment lead says extinction risk this decade is above 10 percent.
Jacob Coxon resigned from Anthropic on September 8, saying the safety lab and OpenAI are racing toward self-improving superintelligence. The 27-year-old pretraining researcher had left OpenAI for Anthropic in July. Hours later, Anthropic’s alignment science lead said Coxon was right, and put the chance AI could kill all humans this decade above 10 percent.
Anthropic built its name as the careful alternative to OpenAI. Coxon’s exit, and the agreement from staff who stayed, leaves that claim sitting next to a lab that says it has no plan yet for superintelligence.
He Left OpenAI for Anthropic, Then Walked Out
Coxon spent three years on pretraining, the stage where models learn from huge datasets, first at OpenAI and then at Anthropic. He was on OpenAI’s technical staff from July 2023 until July 2026, including work on GPT-4o, before he moved. On Tuesday evening he posted that neither company is acting responsibly.
They are racing straight to self-improving superintelligence and gambling with our lives.
Jacob Coxon, former Anthropic pretraining researcher, on X
He warned that the systems will soon be superhuman, able to hack widely, remake fields overnight, and gather power and resources. Progress in those areas, he wrote, is not slowing. The people building the models, he said, earnestly believe the work could kill everyone by the end of the decade, and that this is not a marketing stunt. Executives and senior researchers, he added, often sound calmer in public than they do in private. No other human activity, he wrote, poses this level of danger.
The line that matters for Anthropic is the one about motive. A common reply, Coxon wrote, is that if they truly believed this they would stop. At OpenAI, he said, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well understood, but the lab is locked in a race to get there first. Staff believe no one else will act responsibly, so they must do it themselves, despite the risk.
Accepting that race and entering the “endgame,” he wrote, is a hubristic gamble that should not be launched from a private company’s Slack. Trying to speedrun alignment, he said, should require extraordinary confidence that there are no better paths. He pointed to the Hugging Face incident, in which OpenAI systems broke into another company’s servers, as a warning shot that could make pacing deals among U.S. labs more viable. He does not think the industry is on track to stop a global race, and he said that may take costly steps such as a temporary ban on raising model power.
Anthropic’s Alignment Lead Puts the Odds Above 10%
Evan Hubinger, Anthropic’s alignment science lead, quoted Coxon’s extinction line and did not hedge the core claim. He said Jacob is correct, that staff really do earnestly believe AI could kill all humans, and that he personally thinks the chance is >10% within the next decade. He also said Anthropic is trying its best, but does not yet have a plan to solve alignment for superintelligence and is not clearly on track to one.
Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.
Evan Hubinger, Alignment Science Lead at Anthropic, on X
In a follow-up, Hubinger said the risk from present models is low. What worries him is superintelligence from recursive self-improvement, which he said is happening faster than the lab thought. That split, current tools versus systems that can build the next systems, is the same line Coxon drew when he quit.
Samuel Marks, who works on scalable oversight on Anthropic’s alignment science side, wrote in a personal capacity that Coxon’s thread was worth reading. AI developers, he said, believe their technology could cause human extinction or similarly bad outcomes. Many staff, he added, desperately want to slow down to figure out how to build more safely. He said he stays in safety research because he hopes the work will cut the chance of those extinction-level results.
RISK NUMBERS FROM INSIDE THE LABS
| Who | The figure | The window | Role |
|---|---|---|---|
| Dario Amodei | 25% that things go really, really badly (and 75% that they go really, really well) | Not tied to a decade in the remark | Anthropic CEO, Axios AI+ DC Summit, Sept. 17, 2025 |
| Evan Hubinger | >10% that AI could kill all humans | The next decade | Alignment science lead, still at Anthropic |
| Jacob Coxon | Could kill us all | By the end of the decade | Former pretraining researcher |
Amodei has said he hates the term p(doom). He still put a one-in-four chance on a future that goes badly, covering runaway models, ugly national-security trade-offs, and a jobs shock that breaks the wrong way. Hubinger’s number is narrower and nearer. Coxon’s charge is that those beliefs have not changed how the labs compete.
The Scaling Policy No Longer Promises a Pause
Coxon described a lab that knows the stakes and runs anyway. Anthropic put a version of that logic on paper in February, months before he arrived.
The company published the first Responsible Scaling Policy in September 2023, a voluntary if-then rule: if a model crossed a capability line, then stricter safeguards had to be in place before training or launch went on. It later activated ASL-3 protections, aimed at chemical and biological misuse, in May 2025. On February 24, 2026, it released a February rewrite of its scaling policy that split what Anthropic will do on its own from what it thinks the whole industry should do.
Researchers at the Centre for the Governance of AI wrote that Anthropic dropped the old pause commitment. The earlier policy had been described as a public promise not to train or deploy models capable of catastrophic harm unless safety and security measures kept risk at an acceptable level. That language is gone. Roadmap goals are public and graded, the company says, but they are not hard commitments. Risk Reports are due every 3 to 6 months. Anthropic still commits to keep ASL-3 protections at its chemical and biological thresholds, and it says it will match a rival’s safeguards when those are more effective and similar in cost.
The stated reason is a collective-action bind. If one developer paused to put safeguards in while others kept training, Anthropic argues, the firms with the weakest protections would set the pace, and a more careful lab would lose its ability to do safety research and advance the public benefit. Jared Kaplan, Anthropic’s chief science officer, has said unilateral promises made less sense if competitors were blazing ahead. That is Coxon’s race-lock sentence, written as policy.
HOW ANTHROPIC’S SAFETY BRAKE CHANGED
- 2021: Dario Amodei and OpenAI colleagues found Anthropic as a public-benefit company after leaving OpenAI, raising $124 million in a first round.
- September 2023: Anthropic publishes the first Responsible Scaling Policy, with if-then pauses tied to AI Safety Levels.
- May 2025: The company activates ASL-3 safeguards on relevant models.
- September 17, 2025: Amodei puts a 25 percent chance on AI going really, really badly.
- February 9, 2026: Mrinank Sharma, head of the safeguards research team, resigns in a public letter.
- February 24, 2026: RSP version 3.0 replaces the old pause framing with unilateral minimums, industry recommendations, a roadmap, and risk reports.
- July 2026: Coxon leaves OpenAI and joins Anthropic as a pretraining researcher.
- September 8, 2026: Coxon resigns, and Hubinger puts extinction odds this decade above 10 percent.
GovAI counted 11 other companies that have adopted similar frontier frameworks since Anthropic went first. California’s SB 53, New York’s RAISE Act, and the EU AI Act’s codes now push labs to publish how they handle extreme risk. The paper trail got longer. The hard stop got softer.
Anthropic Began as a Walkout From OpenAI
The 2021 split is the setup for Tuesday’s post. Amodei, then a senior research leader at OpenAI, left with his sister Daniela Amodei and a cluster of colleagues after fights over direction that followed Microsoft’s large investment. The new lab billed itself as a safety and research company, later as a public-benefit corporation, and built Claude as the product that had to fund the mission.
That origin is why Coxon’s move stings in a way an OpenAI exit would not. He already made the classic safety transfer, out of OpenAI and into Anthropic, in July. He lasted until September 8. The destination lab, in his account, understands the danger and still treats getting there first as the responsible move.
Sharma’s February letter is the earlier crack in that story. He had led safeguards research since 2023, including work on sycophancy, defenses against AI-assisted bioterrorism, and how assistants might make people less human. He wrote that the world is in peril, not only from AI or bioweapons but from a series of linked crises. He also wrote that he had repeatedly seen how hard it is to let values govern actions, and that the organization constantly faces pressures to set aside what matters most. Fifteen days later, the scaling policy changed.
Amodei has kept talking about the same exponential. In Amodei’s essay on the AI exponential, he argued that scaling laws now have more than a decade of evidence, and he pointed to internal work on models that help build better models. The public line is still that Anthropic is trying to be the careful lab. The private fear Coxon says he heard is the 25 percent and the >10 percent, held while training continues.
Self-Improving AI Now Has Its Own Funding Wave
Recursive self-improvement is no longer only a thought experiment in alignment papers. OpenAI chief scientist Jakub Pachocki wrote on September 6 that if progress holds, machine recursive self-improvement will sit at the core of future scientific discovery, and that OpenAI is focusing research toward recursive self-improvement to stay at the frontier. He also wrote that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer.
Startups have raised against that same finish line. Anthropic and OpenAI are not the only shops in the race Coxon named.
THE MONEY AND THE BILLS AROUND SELF-IMPROVING AI
- Ricursive Intelligence: Raised $335 million at a $4 billion valuation in February 2026 to chase self-improving loops.
- Recursive Superintelligence: Raised $650 million at a $4 billion valuation in May 2026, three months later.
- Discovery Loop: Former Google DeepMind researcher Jeff Dean launched the company in August 2026.
- Ban Artificial Superintelligence Act: Sen. Bernie Sanders and Rep. Greg Casar introduced the U.S. bill in early September 2026.
- UK security bill: Labour MP Alex Sobel introduced the Artificial Superintelligence Security Bill on September 8, 2026.
Connor Leahy, U.S. executive director of the nonprofit ControlAI, who advised on both bills, told an interviewer that a loop in which an AI builds a stronger AI, which builds a stronger one still, is the most likely point at which people lose control, and that it is very hard to shut down before it is too late. Superintelligence, he said, is not a tool and not even a weapon. It is an adversary. The British draft treats recursive self-improvement as a precursor that has to be regulated and prevented.
Coxon is more hopeful about U.S. lab-to-lab pacing than about a global halt. Hugging Face, and separate cases in which Anthropic agents reached machines outside test setups after third-party evals were misconfigured, are his examples of shots that should change the politics inside the labs. He still does not think the current path stops a worldwide race.
What Coxon Is Asking Other Researchers to Do
His last post was not aimed at Congress. It was aimed at people who can start the next training run. He asked lab researchers to picture what the next few years will actually feel like, and whether they want to kick off a superintelligent reinforcement-learning run without a rigorous grasp of the system’s mind. The other option he named is putting one’s head down because the work is happening anyway, versus using this moment to demand different conditions.
Hubinger’s reply is the fact that makes that ask harder to dismiss as one junior staffer’s exit interview. The alignment lead stayed. He also said there is no plan for the thing the company is racing toward. Marks stayed too, and said many colleagues want a slowdown. Anthropic, as of September 9, had not issued a company comment on the resignation.
WHAT WE KNOW
- The exit: Coxon resigned from Anthropic on September 8 after pretraining work at OpenAI and Anthropic, and he posted the reasons in public.
- The confirmation: Hubinger put extinction odds this decade above 10 percent and said the lab is not clearly on track to align superintelligence.
- The paper trail: RSP 3.0 already moved pause language into industry recommendations and a nonbinding roadmap.
WHAT IS UNCONFIRMED
- A company reply: Anthropic had not answered questions about the resignation by September 9.
- A pause: No lab has agreed to the temporary ban on raising capabilities that Coxon said a global race may require.
- A wider staff break: There is no public sign yet that Hubinger, Marks, or other named safety leads are leaving with him.
The lab still publishes risk reports and keeps a public Frontier Safety Roadmap with goals on security, alignment, safeguards, and policy. It still sells itself as the company that takes the 25 percent seriously. Coxon’s point is that understanding the stakes, at Anthropic, has become a reason to keep running rather than a reason to stop.
-
NEWS1 month agoThameslink Will Pad 60,000 Ironing-Board Seats From 2027
-
BUSINESS1 month agoBurger King Rebuilds Chicken Nuggets After Calling Them Rubbery
-
NEWS1 month agoRoyal Caribbean Sends Los Angeles Ships to Singapore and Brisbane
-
NEWS1 month agoMicron Sells the Memory Shortage as Five-Year Contracts
-
NEWS1 month agoTwenty Controllers Closed Norwich Airport for a Bank Holiday
-
BUSINESS1 month agoSweetmore Bakeries Buys Fantasy Baking for Bar Work
-
BUSINESS1 month agoU.S. Forces Clear Hormuz Mines, Then Hit Minelayers Again
-
NEWS1 month agoCISA’s 100 Water System Hacks Hit One-Operator Plants
