8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

Daily Podcast full article

Anthropic pushes Claude into labs

Anthropic says Claude has produced a complete Lean-verified formalization of Fermat’s Last Theorem, turning one of mathematics’ most famous proofs into a machine-checkable artifact and raising the stakes for AI-assisted science.

Generated September 5, 2026 at 12:35 AM UTC1337 words
AI-generated illustration

The claim

Anthropic has moved Claude deeper into the research lab with a claim that would have sounded implausible only recently: the company says its model produced the first complete computer-checked proof of Fermat’s Last Theorem in Lean, the formal proof assistant used to verify mathematical reasoning step by step . The work, announced on September 4, 2026, is not presented as a brand-new route to Fermat’s Last Theorem; Anthropic says Claude formalized a proof path following the modern Wiles and Taylor-Wiles tradition, rather than discovering a wholly independent proof .

That distinction matters. Fermat’s Last Theorem states that no positive integers a, b and c satisfy aⁿ + bⁿ = cⁿ for any exponent n greater than 2 . The theorem was famously proved by Andrew Wiles in the 1990s, after a first announced proof exposed a gap and required further work with Richard Taylor before publication . Anthropic’s claim is therefore not that Claude has solved an unsolved problem, but that Claude has translated a vast amount of difficult mathematics into a form that Lean can check mechanically .

The scale is the startling part. Anthropic says Claude worked largely autonomously for 11 days, generated 13 million lines of Lean code, and proved 29,500 intermediate theorems used in the final proof . TechNewsReel separately reported the same core figures, describing the work as a first computer-checked proof of Fermat’s Last Theorem produced by Claude over an 11-day period . CryptoBriefing also reported the 11-day formalization and said the process had previously been expected to take years .

What Lean changes

Lean is not a referee in the ordinary human sense. It does not read a mathematical paper, decide whether an argument is elegant, or judge whether an exposition is illuminating. It checks whether formal statements follow from prior formal statements under the rules encoded in the system. In that environment, every hidden assumption has to be made explicit, and every skipped step has to be supplied.

That is why the Anthropic announcement is potentially important. A normal proof in number theory can be accepted after months or years of expert reading, with trust distributed across known results, specialists, journals and reputation. A Lean proof, by contrast, aims to reduce the question to whether a formal object type-checks. Anthropic says the finished proof was checked by Lean and used only Lean’s three standard axioms, while a comparator confirmed that the theorem statement matched Mathlib’s own Fermat’s Last Theorem statement .

The formalization burden has been the bottleneck. Human mathematicians write for other humans; they suppress routine algebra, cite major theorems by name and rely on shared context. A proof assistant demands a far more granular reconstruction. Anthropic says the community expected a full Fermat formalization to take years, and notes that the blueprint used for the initial phase of the project runs to 86 pages . If Claude really compressed a substantial part of that workflow into 11 days, the achievement is less about “AI knows Fermat” and more about “AI can operate inside a formal laboratory at scale.”

How Claude was organized

Anthropic says the project was led through the work of Tianyi Peng, an Anthropic researcher whose Columbia University group builds tools for AI formalization . According to the company, early attempts failed because agents lost track of the state of the proof and stopped collaborating effectively . The successful run used Prove2Me, a collaborative platform designed to manage formalization tasks through a directed graph of theorem statements .

That architecture is a key part of the story. Large mathematical projects are not single prompts; they are dependency networks. Definitions must be established before lemmas, lemmas before propositions, propositions before the final theorem. Anthropic says Prove2Me helped Claude agents decide what to prove next, compile work more efficiently, and search for reusable theorem statements through natural-language descriptions . In other words, Claude did not merely emit a long answer. It operated as a swarm of agents moving through a structured research plan.

The company also says the system consumed about six billion output tokens using a general-purpose internal research model roughly comparable to Claude Fable 5.1 . That number is a reminder that the breakthrough, if it survives scrutiny, is not a cheap party trick. It required scaffolding, compute, orchestration, a proof assistant, and a carefully managed corpus of intermediate goals.

The role of Kevin Buzzard

Anthropic says it shared the resulting proof with Kevin Buzzard, the Imperial College London mathematician closely associated with efforts to formalize Fermat’s Last Theorem in Lean . In Anthropic’s account, Buzzard characterized the work as an “extraordinary autoformalization achievement” and said it proved Fermat’s Last Theorem with no assumptions other than the axioms of mathematics .

That endorsement is important but should be read precisely. It is not the same as years of broad community inspection. The current public status is that Anthropic has announced the proof, reported Lean verification, and cited review by a leading formalization expert . The next phase is outside examination: mathematicians and Lean specialists will want to inspect the artifact, reproduce the checks, test the assumptions, and understand how much of the proof is readable, reusable and maintainable.

This is especially important because formal artifacts can be correct in the kernel’s sense while still being difficult for humans to audit conceptually. A 13-million-line proof is a scientific object as well as a software object. Its mathematical value will depend not only on whether it checks, but also on whether it can be navigated, simplified and incorporated into the wider ecosystem of formal mathematics.

Why it matters beyond Fermat

The broader implication is not that every major theorem is suddenly solved. Fermat’s Last Theorem was already known to be true. The significance is that a frontier model may have helped convert a major human proof into a durable, machine-checkable object at unprecedented speed . That is a different kind of scientific contribution: not discovery alone, but verification infrastructure.

Anthropic frames the achievement as a way to reduce the burden on human referees and to help mathematicians cope with an era in which AI systems may generate more claimed proofs than humans can comfortably check . That argument is persuasive in one respect. If AI accelerates conjecture-making and proof-writing, verification becomes the scarce resource. Formal proof assistants offer a way to keep the pace of trust closer to the pace of generation.

Still, caution is warranted. Formal verification does not replace human understanding, and Anthropic itself says a formalized proof should not replace a human-readable exposition . A Lean certificate can tell the field that a statement follows from specified foundations; it cannot by itself explain why the proof is beautiful, how the ideas relate to a wider theory, or which parts should guide future research.

A new laboratory standard

The current state of the story is therefore extraordinary but bounded. Anthropic says Claude has produced a Lean-checked formalization of Fermat’s Last Theorem, using a multi-agent workflow, Prove2Me, 13 million lines of code and tens of thousands of intermediate theorems . Independent coverage has repeated the central claim and framed it as a milestone for AI in formal mathematics . The open question is how the broader mathematical community will evaluate, reproduce and absorb the artifact.

If the proof holds up under outside scrutiny, the milestone will mark a shift in what “AI for science” means. The frontier would no longer be limited to fluent explanations, code generation or plausible research drafts. It would include models working inside formal systems, producing outputs that machines can verify line by line. That is why Anthropic’s move matters: Claude is not just being pushed into chat windows or coding tools. It is being pushed into the lab, where claims must survive the harsh discipline of formal proof.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]Formalizing Fermat's Last TheoremSep 4, 2026, 12:00 AM UTC
  2. [2]Anthropic's Claude Completes First Computer-Checked Proof of Fermat's Last TheoremSep 4, 2026, 12:00 AM UTC
  3. [3]Anthropic’s Claude formalizes Fermat’s Last Theorem in 11 daysSep 4, 2026, 12:00 AM UTC

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.