About two weeks before OpenAI's agents produced their Navier–Stokes result, Noam Brown was discussing with a researcher at a frontier lab how long it would take AI to solve a Millennium Prize problem. The researcher was willing to bet Brown $1,000 that it would not happen until after 2027; he expected about 2030. Brown took the bet. Even so, Brown said, he too had expected it to take longer than it did.
Brown, an OpenAI researcher, told the story on the Dwarkesh Podcast in an episode published on September 17, 2026. He and the show's host were discussing what the recent burst of AI progress in mathematics says about recursive self-improvement, or RSI. The term describes AI systems doing the research that produces better AI systems, which could then improve research further. The host argued that the maths results preview automated AI research. Brown agreed about the opportunity but not about the speed. He expects a large acceleration, perhaps around threefold. He does not expect an overnight "intelligence explosion", because AI research depends on experiments, and those take time and computing power.
A ladder of tenfold steps
OpenAI's announcement of September 8 describes about 10,000 AI agents working at the same time. According to the company, they reached an answer about 88 hours after launch, and a formal check in the proof language Lean took another 17 hours. The company says the agents found a finite-time singularity: a fluid flow that starts smooth but whose velocity becomes unbounded. That claim addresses the versions of the problem in the official prize formulation that allow a smooth outside force to act on the fluid. It does not show that solutions always stay smooth.
Brown explained why the result came earlier than he expected. He had been tracking AI progress in maths by how long each type of problem would take a capable human. GSM8K grade-school word problems take a mathematician about five seconds. Problems from the MATH benchmark of competition questions take an expert maybe a minute. Problems from AIME, a qualifying exam for the US Mathematical Olympiad team, take a good mathematician about ten minutes. Brown said models mastered each level about a year after the one before. These are his rough estimates, but they describe a roughly tenfold increase each year. OpenAI's o1 report from September 2024 gives one marker on this path: it says o1 averaged 11.1 of 15 AIME problems on a single attempt, compared with 1.8 for GPT-4o.
On that trend, Brown said, a gold medal at the International Mathematical Olympiad in 2025 was "very sensible": an olympiad problem takes about 100 minutes. The next step would be about 15 hours of human work, which he judged far too little for a Millennium Prize problem. His forecast was that it would not happen in 2026, probably not in 2027, and maybe in 2028. "So it did happen a lot faster than I expected," he said.
Brown said even the 2025 gold surprised people. He recalled that even people at OpenAI had thought it almost impossible for a general-purpose language model with no tools or internet access to win gold.
Solving problems versus choosing them
The host, who described himself as a podcaster reasoning about the field as "a total outsider", traced the same rise. In 2024, AI solved a few high-school competition problems. In 2025, it won olympiad gold. Earlier in 2026, it solved open Erdős problems, though perhaps ones people had not tried very hard on, or whose solutions resembled something already in the literature. Now it had solved a Millennium Prize problem, for which, he said, there was "no story of why this should have been easy."
He acknowledged an objection raised by the mathematician Terence Tao and by Toby Ord. In a recent essay, Ord argues that mathematics is more than proving statements. It also means choosing questions worth asking and inventing new concepts or whole fields. Ord calls the evidence that AI can do those things sparse, partly because such skills lack clear benchmarks and easy-to-check rewards. He also allows that AI might acquire them quickly.
The host's answer was that this gap matters less in machine learning. There, researchers mostly want to solve well-defined problems, such as making models learn from less data. By that standard, the kind of progress now arriving in mathematics looked to him structurally similar to what would directly speed up AI research.
Brown agreed with much of this. Today's models are "jagged", he said: brilliant in some ways and weaker than human mathematicians in others. They are not very good at posing new problems or judging which branches of mathematics are worth developing. He would be "thrilled" if AI stayed a complement to human mathematicians. But he expects models to improve across the board over time, and how long that takes depends on how long the tail of things they are bad at turns out to be. He also agreed that their particular strengths suit RSI. AI research has clear metrics, he said: "if you can make it do better on those metrics, then you've succeeded."
Thinking is not the only bottleneck
The host offered what he called an intuition pump. Agents might spend more cognitive effort in a week on a long-standing machine-learning problem than the whole field has spent on it so far. Experiments need compute, he granted. But by his estimate, OpenAI will have enough by the end of next year to give each of 10,000 agents the capacity to run an experiment the size of GPT-3 every day. For scale, the GPT-3 paper describes a largest model of 175 billion parameters trained on 300 billion tokens.
Brown called the intuition pump "pretty accurate" and said the models' uneven strengths are probably especially useful for RSI. But one difference matters. Mathematics is mostly "bottlenecked by thinking really hard," which models do well. "When you look at things like RSI, you do have to run experiments."
He offered a thought experiment. Put all the most brilliant people in the world at OpenAI, but give them 100 times less computing power. Would they make more progress than today's staff with today's compute? Brown suspected they would make less, "a lot less." Asked whether it would be 100 times less, he said no. In his view, intelligence without compute has limits. Experiments often have to run one after another, because each depends on training a model or waiting for the last result. And every experiment needs GPUs, the specialized chips that AI systems run on.
That is why Brown expects "a significant speed up" but not a jump to going 100 times faster. "If that exponential is like three X faster, that is massive," he said. Pressed for a number, he said he could see things going three times faster. He compared that to making the progress of the past three years in one year. The host put it another way: it would be like going from non-reasoning models, before o1, to GPT-6 Astra in a single year.
Brown did not treat three times as a firm prediction. He said he could be "totally wrong". Progress might speed up by only 50%. A tenfold speed-up was unlikely but possible, and he did not rule out an overnight explosion. "There's a lot of uncertainty here," he said.
A curriculum that could run dry
Brown also named one way language models could fail to follow the path of game-playing AI. The host had said it was striking that models presumably trained on much easier, checkable problems could tackle something as ambitious as a Millennium Prize problem. Brown replied that OpenAI does train on very hard problems. But he said the problems themselves could become a limit.
DeepMind's AlphaZero learned games by playing against itself, so its opponent improved along with it. Brown called that "an infinite curriculum". He recalled that Go-playing systems went, within about a year, from beating a European champion to beating the world champion to being far beyond any human. Language models trained with reinforcement learning, a method that rewards successful answers, instead need to be given problems. "If the problem is so easy that I could just solve it in a second, it's not really learning anything," Brown said. If labs run out of problems hard enough to challenge the models, progress could become much harder. He added that this had not yet become a wall, and that he thought there would be ways around it.
Hard to measure, even from inside
The host asked when Brown expected AI research work to be 95% automated. Brown said he could not say, partly because the acceleration is hard to measure. He pointed to OpenAI's recent report on internal research acceleration. The report says that by mid-August, median daily coding-agent use exceeded $600 at API prices, while the 90th percentile exceeded $7,000. It presents these as usage measures, not measured productivity multipliers. It notes that experiments per active experimenter rose but available compute also grew, and that more than half of successful four-to-eight-hour agent tasks involved human intervention.
Brown explained why usage is hard to convert into a speed-up. When a human directs AI agents, it is unclear how much of the work to credit to each. Models are also especially good at certain tasks, such as checking every data point in a dataset for quality. Researchers therefore use them heavily for that work, which can become 100 times faster and better while other work changes little. And anything that suddenly becomes 100 times faster and better gets done much more often. Brown said this leaves two different questions. How much faster can a lab now do the work it did three years ago? And how much slower would today's work have been three years ago? He was confident of one thing: work is going faster than even a year ago because of AI, and he expects that acceleration to continue.
Forecasts on a shrinking horizon
The discussion also turned to what the current pace of progress alone would imply. Even if progress did not speed up and simply continued as it is, the argument went, a given amount of compute could run an effective population of AI workers three times larger each year, while compute itself keeps growing. On that extrapolation, each lab could have enough compute by the end of 2030, or sooner, to run hundreds of millions of human-level intelligences, and by the mid-2030s or earlier the equivalent of many Earths' worth of them. The extrapolation was offered as a base case, with an explicit admission of not knowing.
Brown agreed that progress is "really fast" but declined to put a date on full automation. "I don't know what the world looks like in 2030. That's the truth," he said. He described someone who worked on the Navier–Stokes effort. That person used to feel comfortable predicting AI progress 12 months ahead; now he is not comfortable going beyond three months. Brown said views were shifting inside OpenAI too. People who expected things to take longer are starting to feel that progress is faster than expected, and he called that "an increasingly common conversation to have."