Claude Just Formalized a Proof of Fermat's Last Theorem
Anthropic says its Claude model produced the first end-to-end, computer-checked proof of Fermat's Last Theorem โ a task mathematicians expected to take years โ in just 11 days, working largely autonomously. Here's what that actually means, and why it's a bigger deal than it might sound.
โก Quick facts
- Time taken: 11 days, largely autonomous
- Scale: 13 million lines of Lean code, 30,300 theorems proved (29,500 used in the final proof)
- Compute used: Roughly 6 billion output tokens
- Tool used: Prove2Me, an open-source platform from Anthropic researcher Tianyi Peng's group at Columbia University
What Claude actually did (and didn't do)
It's important to be precise here: Claude did not discover that Fermat's Last Theorem is true. Andrew Wiles proved that in 1994, closing a problem that had stood unsolved for 358 years. What Claude did was formalize that proof โ convert Wiles' mathematical reasoning into Lean, a programming language built for writing proofs a computer can check step-by-step with total rigor, leaving no room for a subtle human error to slip through unnoticed.
Why formalization is its own enormous task
Wiles' original proof is famously dense, drawing on advanced areas of number theory. Checking that a proof this complex is airtight has traditionally required teams of expert mathematicians combing through it by hand โ a process that can take years, precisely because human review can miss subtle gaps. Formalizing it in Lean removes that uncertainty: once Lean accepts a proof, it's verified using nothing but its three standard axioms โ no trust in any individual mathematician required.
The scale of what Claude produced
Over 11 days, Claude generated 13 million lines of Lean code and proved 30,300 individual theorems, of which 29,500 were ultimately used in the completed proof โ meaning it had to build almost the entire mathematical scaffolding underneath Wiles' argument from the ground up, in a form Lean could verify. The run consumed roughly 6 billion output tokens in the process.
What made it possible: Prove2Me
The breakthrough came from giving Claude access to Prove2Me, an open-source platform built by Anthropic researcher Tianyi Peng's group at Columbia University. It's specifically designed to support the kind of large-scale, long-horizon formalization work this project required โ the same broad trend toward AI agents sustaining complex, multi-step tasks over long stretches of time this site has covered with other recent model releases.
Why this matters
Mathematicians expected formalizing Wiles' proof to take years of dedicated human effort. Claude did it in 11 days. That's a meaningful signal about how far AI-assisted formal reasoning has come โ not just generating plausible-sounding text, but producing rigorous, machine-verifiable mathematics at a scale and speed no human team could match. It's also a preview of a future where formally verifying complex proofs, software, or even legal and financial reasoning could become dramatically faster and more trustworthy.
Frequently asked questions
What did Claude actually do with Fermat's Last Theorem?
It formalized Andrew Wiles' 1994 proof in Lean, a language a computer can rigorously verify -- it didn't prove the theorem from scratch.
How long did it take, and how big was it?
About 11 days, generating 13 million lines of Lean code and proving 30,300 theorems using roughly 6 billion output tokens.
What is Prove2Me?
An open-source platform from Anthropic researcher Tianyi Peng's group at Columbia University that enabled this large-scale formalization work.
Why does formalizing a proof matter?
It removes the years of human effort otherwise needed to verify a proof is airtight -- Lean confirmed this one using only its three standard axioms.
Did Claude discover Fermat's Last Theorem is true?
No -- Andrew Wiles proved that in 1994. Claude's achievement was formalizing that existing proof for computer verification.