25 September 2026
Heard In AI

A story we follow

Anthropic Formalizes Fermat's Last Theorem

Tracks Anthropic's AI-assisted formalization of Wiles's proof and lessons for coordinating large mathematical agent projects.

A story page follows one specific event across podcast discussions, with an overview and a timeline of what changed. It updates when new episodes discuss the event, and you can get those updates by email or push. How our formats work

Overview

Anthropic reported an AI-assisted Lean formalization of an existing proof of Fermat's Last Theorem, rather than discovery of a new proof. Moonshots highlighted the roughly 13-million-line artifact and thousands of intermediate theorems, and treated it as evidence that large agent workflows need explicit objectives and mechanisms to consolidate sprawling output. In a later Moonshots episode, Robinhood CEO and Harmonic founder Vlad Tenev used the formalization as a model for verifiable AI output: no human will read a 13-million-line proof, but a person can check the one-line theorem statement and let Lean check the rest, which he estimated saves well over 90 percent of the effort. He predicted AI-generated code will come with certificates of correctness. Alex Wissner-Gross objected that Fermat is an easy case because its statement is simple, while specifications such as safety in an agentic environment are hard to state, and a strong model could slip in definitions that help it and harm humans. Tenev said models are becoming more faithful when formalizing statements. Tenev's retelling of the history, that repairing Wiles's error took six or seven years, differs from Anthropic's own report, which puts the repair at roughly a year.

What changed

Dates show when each podcast discussion was published.

  1. Interpretation

    Vlad Tenev presented the Fermat formalization as a template for certifying AI output: humans check the short theorem statement and Lean checks the 13 million lines. Alex Wissner-Gross countered that this works only when the statement is simple, and questioned whether the method can extend to agent safety. Tenev's account of how long Wiles's repair took differs from Anthropic's history.

  2. New information

    The hosts distinguish formalizing Wiles's proof from proving the theorem for the first time and draw practical lessons about coordinating agents around a concrete final artifact.

Podcast discussions

Sources

  1. 01
  2. 02
  3. 03
  4. 04

Our coverage

How agent teams turned Fermat's proof into 13 million checked lines

On Moonshots with Peter Diamandis, a panelist interrupted an argument about AI regulation to read a headline off his feed: Anthropic had formalized Fermat's Last Theorem. Anthropic's report describes dozens of agents working eleven days, about six billion output tokens and 30,300 intermediate theorems — plus a piece of bookkeeping software that stopped runs from losing track of their own work. The panel's takeaway was about how to narrow enormous machine output into one result you can build on.

· Updated 5 min read

Version history

  • 25 Sep 2026 · Version 2

    Vlad Tenev presented the Fermat formalization as a template for certifying AI output: humans check the short theorem statement and Lean checks the 13 million lines. Alex Wissner-Gross countered that this works only when the statement is simple, and questioned whether the method can extend to agent safety. Tenev's account of how long Wiles's repair took differs from Anthropic's history.