8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesFor youTopicsVideosYT channelsArchivesSearchFavorites

Daily Podcast full article

Anthropic expert puts AI extinction risk above 10%

An Anthropic alignment researcher has publicly put a personal probability of more than 10% on AI killing all humans within the next decade, after colleague Jacob Coxon resigned and accused leading labs of racing toward self-improving systems they may not be able to control.

Generated September 9, 2026 at 10:40 AM UTC1530 words

A warning from inside the lab

The new alarm over artificial intelligence risk is not coming from an outside critic, a political campaigner or a science-fiction writer. It is coming from inside Anthropic, one of the frontier AI companies building the systems at the center of the debate.

Evan Hubinger, identified in recent reports as Anthropic’s Alignment Science Lead, wrote on X that “AI could kill all humans” and gave his own estimate as greater than 10% within the next decade . His post followed the resignation of Jacob Coxon, an Anthropic researcher who said he was leaving the company and the AI industry after three years of pretraining work at OpenAI and Anthropic .

The number is not a formal Anthropic forecast, and it should not be read as a measured probability in the way a weather model estimates tomorrow’s chance of rain. It is a subjective judgment about an unprecedented technological risk. But the source matters. When someone working on alignment at a leading lab says, publicly and in his own name, that the chance of human extinction from AI is above one in ten, the claim becomes a governance event as well as a philosophical one.

Coxon’s resignation gave the warning its immediate trigger. He said neither OpenAI nor Anthropic was acting responsibly and accused them of “racing straight to self-improving superintelligence” while “gambling with our lives” . In separate reporting, Coxon said people building AI “earnestly believe” it could kill everyone by the end of the decade, and that senior people often express such fears more bluntly in private than in public .

What Hubinger did — and did not — say

Hubinger’s statement is striking because it combines three elements that are usually kept separate: a catastrophic outcome, a near-term time horizon and an insider’s assessment of institutional readiness.

According to Forbes, Hubinger wrote that he personally sees a greater than 10% chance of AI killing all humans within the next decade, while also saying that Anthropic is trying its best but does not yet have a plan to solve alignment for superintelligence and is not clearly on track to do so . ABC reported the same core sequence: Coxon quit, Hubinger replied that Coxon was correct, and Hubinger added that Anthropic did not yet have a plan for superintelligent alignment .

That does not mean Hubinger said current public AI systems are about to exterminate humanity. Forbes reported that he later pointed to Anthropic’s risk framing and clarified that he was worried about superintelligence arising from recursive self-improvement rather than present-day models as they exist now . In other words, the fear is not that today’s chatbot suddenly becomes Skynet. The fear is that increasingly capable systems could help improve their successors, accelerate the pace of AI development and eventually produce systems that humans cannot reliably understand, constrain or shut down.

Axios framed the same dilemma as a race dynamic: companies and researchers worry that slowing down could mean falling behind, while continuing to race could mean losing control . That is the governance trap at the center of this story. If every lab believes a rival will keep going, then even sincere safety concerns may not translate into restraint.

Why “more than 10%” lands differently

Risk professionals do not need certainty before taking a catastrophic hazard seriously. A bridge, reactor, aircraft or medical device would not be cleared on the logic that a one-in-ten chance of total disaster is tolerable. The reason Hubinger’s number matters is therefore not that it can be independently verified. It cannot. The relevant point is that a technically informed insider at a company building frontier AI considers the probability seriously non-negligible.

The Next Web argued that the estimate should be treated as a personal probability, not as a company prediction or an empirical measurement . That distinction is essential. There is no historical data set of superintelligent AI takeovers from which to derive a frequency. Assessments of existential AI risk mix technical inference, judgment about institutional behavior, beliefs about future capability growth and assumptions about whether alignment techniques will scale.

Still, subjective probabilities guide real policy decisions all the time. Governments use them in intelligence analysis; companies use them in cybersecurity; public health officials use them when preparing for low-frequency, high-consequence events. A 10% extinction-risk estimate from a frontier-lab alignment lead is therefore not a statistic to be accepted blindly, but neither is it a remark to wave away as hyperbole.

The debate is made harder by the incentives surrounding AI. Axios noted that critics see a possible strategic upside for major labs in emphasizing catastrophic risk: it can raise the political salience of the technology, attract attention, and support regulation that entrenches large incumbents . But Axios also reported that people inside leading AI companies have sounded increasingly worried in recent months, in part because they see unreleased systems and internal trajectories the public cannot evaluate .

Both things can be true. Companies can have incentives to dramatize risk, and their researchers can genuinely believe the danger is real. That is precisely why outside oversight matters.

The resignation as a governance signal

Coxon’s departure is important because he did not merely complain about one model release or one internal decision. He described a structural problem: a competitive race toward systems that may be able to improve themselves.

Reports describe Coxon as a pretraining researcher, meaning he worked on the computationally intensive phase in which large models absorb vast quantities of data before later tuning and deployment . That places him close to the capability-building side of AI, not only the policy or communications side. He said he had worked at both OpenAI and Anthropic, giving his warning added weight because it concerns the behavior of more than one leading lab .

Coxon also called for different conditions for researchers and pointed to pacing agreements or temporary limits on capability advances as possible responses . That emphasis matters. The story is not just “one researcher quit.” It is a warning that individual conscience may not be enough if the institutional structure rewards speed.

Axios reported that another Anthropic figure, Samuel Marks, who leads scalable oversight, also said AI developers believe their technology could cause human extinction or similarly bad outcomes, potentially within the next few years . That makes the episode broader than a single resignation and a single reply. It suggests that the language of catastrophic risk is now being spoken openly by multiple people inside the same frontier lab.

What regulators will hear

For regulators, the useful takeaway is not that one X post proves a specific extinction probability. It does not. The takeaway is that frontier labs themselves are producing testimony that the upper tail of AI risk may be extreme.

ABC reported that more than 1,300 employees from frontier AI companies had signed a July statement warning of a real risk that capability development could accelerate beyond society’s ability to understand or control resulting systems . The same report said the statement called for international tools to deliberately pace advanced AI development . In that context, Hubinger’s estimate becomes part of a wider pattern: insiders are not merely asking for better product testing, but for mechanisms that can slow or coordinate the frontier itself.

That is politically explosive. If AI risk is treated as ordinary software risk, then transparency reports, red-team testing and post-deployment monitoring might seem sufficient. If it is treated as a possible extinction risk, then policymakers may look at compute licensing, model evaluations before release, incident reporting, liability, cross-border agreements and emergency shutdown powers.

The hard part is designing rules that reduce catastrophic risk without simply giving the largest companies a protected moat. The same labs warning of danger are also competing for talent, capital, customers and strategic position. Any regulatory response has to account for that conflict.

The line Anthropic now has to walk

Anthropic has built much of its public identity around safety. That makes Hubinger’s candor both consistent with the company’s brand and uncomfortable for it. A lab that says the danger is real invites trust for taking the issue seriously, but it also invites scrutiny over why it continues to build more powerful systems without, by its own insider’s account, a clear path to aligning superintelligence.

The most important question is not whether Hubinger’s personal number is exactly right. It is whether Anthropic, OpenAI and other frontier labs can demonstrate that their governance mechanisms are stronger than the race incentives Coxon denounced.

If the probability of catastrophe were 0.1%, regulators would still have reason to ask hard questions. At greater than 10%, even as a subjective estimate, the burden of proof shifts. The public does not need to accept every apocalyptic scenario to demand controls proportionate to the downside.

Skynet is still fiction. But the people building the nearest real-world analogues are now publicly debating the odds of losing control. That alone is no longer a plot device. It is a policy problem.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]Anthropic Alignment Lead Warns AI Could ‘Kill All Humans’ As Researcher QuitsSep 9, 2026, 4:14 AM UTC
  2. [2]Anthropic insiders warn AI could kill all humansSep 9, 2026, 10:04 AM UTC
  3. [3]Anthropic researcher quits saying AI developers believe 'it could kill us all'Sep 9, 2026, 9:11 AM UTC
  4. [4]An Anthropic researcher quit saying AI labs are gambling with our livesSep 9, 2026, 7:33 AM UTC
  5. [5]‘Gambling with our lives’: Anthropic researcher quits over AI labs’ ‘irresponsible’ raceSep 8, 2026, 9:00 PM UTC

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.