One panelist on The Diary of a CEO raised the question behind the whole debate. If AI systems are solving some of the most famous open problems in mathematics, "how much harder is it to have an AI solve the problem of make me a smarter AI, make me AI architectures that learn faster?"
The first answers were short: "Possibly quite a lot." "Could be a lot." "I hope it's a lot."
The episode was published on September 17, 2026, and had four guests. Ed Zitron is a tech critic and CEO of EZPR, a technology and business PR and research agency. Andrew McAfee is a principal research scientist at MIT and co-director of its Initiative on the Digital Economy. Nate Soares is president of the Machine Intelligence Research Institute and author of If Anyone Builds It, Everyone Dies. Roman Yampolskiy is a computer scientist who works on AI safety, cybersecurity and digital forensics.
The math result prompted the argument. The dispute was over a different step: whether AI systems will soon improve the research process that produces their own successors.
The loop Yampolskiy is worried about
When asked to make his case, Yampolskiy began by splitting "AI" into three technologies. The first is narrow tools that make people more productive. As a computer scientist and engineer, he said, he wants more of them: "We know how to control them, how to make them safe." The second is the kind of system "we're starting to have now," which he called "GPT-6 level, human level, AGI level." In his view, these systems are "unsafe like a human would be unsafe."
His worry is the third: what happens "if we introduce them into the research cycle" as automated scientists and engineers. Right now, he explained, humans do the research that produces the next model, with AI tools taking on more of the programming. "What if the whole process is fully automated? What if GPT-6 is writing GPT-7?"
Asked whether this was what people call recursive self-improvement, Yampolskiy said it was. The term means AI systems improving the methods used to build better AI, which then improve those methods further. Someone in the conversation said that outcome is "not a foregone conclusion." The reply was that the top labs predict they will get there, "introducing junior machine learning researcher in 2026" and wanting the cycle to start in 2027. Once it starts, the argument went, the result would be superintelligence, "a system smarter than all of us at everything."
The conversation then turned to "fast takeoff," a phrase one speaker said they had heard from Sam Altman and others. One panelist explained the dispute behind it. Some people expect automated research to take years anyway, because physical experiments still have to be run. Fast takeoff means the process takes "a month, a week, a day, a second" instead of a year. The panelist described it with an image: "10,000 agents, each one smarter than all of us, doing research 24-7."
What OpenAI actually reported
The Millennium Prize Problems are famous mathematical questions announced in 2000, each with a $1 million prize. One concerns the Navier–Stokes equations, which describe how fluids such as water and air move. On September 8, OpenAI announced a proof about these equations. It said that a three-dimensional fluid that starts out smooth, with smooth outside forcing and finite energy, can develop a singularity in finite time. A singularity here is a point where the equations stop giving sensible answers. The company described a vortex that becomes more and more concentrated until its speed grows without limit. OpenAI also said the proof had been formalized in Lean, software that checks each step of a proof mechanically. On September 11, the Clay Mathematics Institute said the problem appears to have been settled. It added that evaluation and credit would follow its prize rules through a deliberately unhurried process.
One panelist described the result as "a swarm of 10,000 open AI agents running for 11 days." OpenAI's own account differs on the timing and on the role of people. It says the successful group used about 10,000 agents running at the same time, and that human researchers chose the problems. After one intermediate result on the related Euler equations, the humans shifted resources. They also pulled together the agents' partial findings. According to OpenAI, the solution arrived about 88 hours after launch, and formalization and checking took another 17 hours.
Panelists also questioned where the result came from. One said OpenAI had been racing human researchers who were close to solving the problem themselves, and that "it's unclear how much of their work" OpenAI used. Another said there had been no confirmation of whether OpenAI had been drawing on two scientists who were using LLMs to solve the problem. In a September 10 update, OpenAI said its investigation had ruled out any influence from mathematician Tristan Buckmaster's recent prompts to its Codex coding tool. The company distinguished its result from related work by Buckmaster and Levent Alpöge that was happening at the same time. It also said it would not claim the Millennium Prize.
Panelists also mentioned claims about other Millennium problems. One said two others had been claimed. Another said they had seen rumors that several had been claimed but had not looked into everything closely. In the discussion, those claims remained rumors.
Math result versus self-improvement
One panelist said the other side was making "a logical leap." In that panelist's view, it was "really interesting" to see language models help with a result like this. But it was different from a case where "AI did this completely on its own." The panelist agreed that such a case would need to be contained, understood and prepared for, "or indeed slow down until we understand what that means." A skeptical speaker who said "I'm not a scientist" drew a similar line. In that speaker's view, the fluid problem is "a very specific mathematical scientific principle," while building better AI is a more general problem that could go many different ways.
The replies came from the risk side. One speaker argued that even if the AIs had been trained on human work, they "did go a bit further," and that many humans are also doing AI research. Another recalled what happened when AI systems began solving gold-medal problems from the International Mathematical Olympiad, the most prestigious math competition for teenagers. Critics called them "just problems for kids" and said: "Wake me up when the AIs can solve millennium problems." The speaker went on: "Now the AIs are solving millennium problems. And like, where are the people waking up?" The same speaker said they hoped the systems were "cheating off of people's notes." But "a year ago, if you said millennium problems don't take that much creative thinking, you would have been laughed out of the room."
One risk-focused panelist gave a specific estimate. Six months from now, could a lab put 100,000 agents to work for 12 days on designing a smarter AI architecture and have it work? "I think more likely than not, they won't be able to do that yet," the panelist said, "but I think 10% chance maybe that if they try that in six months, it works." The panelist added: "I'm not saying that they will be able to make smarter AIs in six months. I'm saying six months ago, millennium problems looked like they were out of reach."
A separate argument concerned large language models (LLMs), the technology behind today's chatbots. One risk-focused panelist said they had been "really hoping that the LLMs will run out of steam and they keep on not running out of steam." In that panelist's view, even a plateau might not end the concern. The question is whether LLMs stall "at a point where they can do automated AI research and find some other architecture that's better than LLMs." That would be a new basic design for AI, possibly a cheaper and more efficient one.
What partial automation looks like
Two days before the math announcement, OpenAI published a separate picture of AI-assisted research. Its September 6 internal analysis says coding agents took on more research work between January and August, including fixing infrastructure problems and monitoring experiments. High-level planning remained a small share of what the agents produced. Among successful tasks that OpenAI estimated would take a person four to eight hours, more than half still needed human intervention. That analysis excluded uncertain outcomes and small categories. The company also warned that lines of code and numbers of experiments are imperfect measures of research progress. Because OpenAI had more computing power over the same period, it said, gains are hard to credit to the agents alone.
The argument over AI 2027
The dispute over timing centered on AI 2027, a scenario by Daniel Kokotajlo and his colleagues at the AI Futures Project. Its milestones were read out during the exchange. Superhuman AI coders arrive in March 2027, followed by a superhuman AI researcher in August and a superintelligent AI researcher in November. At that point, AI progress runs 250 times faster than human-only research. The scenario reaches artificial superintelligence in December 2027.
A skeptical panelist said the scenario already depended on AI "teaching itself": "Without that link, AI 2027 kind of falls apart." The same panelist agreed that regulation is needed. But in that panelist's view, focusing on AI 2027 "gets away from actually fixing the problem" by shifting attention to the future and away from what to do today. The panel also went through the scenario's predictions for 2026. They included huge increases in computing and power, AI agents becoming normal, the rise of coding agents, AI systems faking alignment and deceiving people, and industrial espionage. One panelist said the authors had "nailed those predictions better than me."
The authors' takeoff forecast, by Kokotajlo and Eli Lifland, is more conditional than the month-by-month story. It starts by assuming that a superhuman coder arrives in March 2027. It then estimates how long later advances would take with humans alone and applies simulated speedups from automating AI research, holding constant the computing power used for training. The August, November and December dates belong to a narrative about labs racing each other. Starting from the March 2027 assumption, the forecast puts its median date for general superintelligence at April 2028. Its 80% range runs from June 2027 to beyond 2100. A December 2025 note on the page stresses that the forecast depends on intuitive judgment.
Asked about timelines, one risk-focused panelist said that if recursive self-improvement starts this year, "2027 looks as reasonable as any other year" for AI to go beyond human level. When the date was challenged, a panelist answered: "Fine, 30, 35. Does it make a difference? We are gambling all of humanity."
Near the end of the exchange, one panelist said, "I think we cannot rule out this scenario." Another replied, "I think we can't rule it in." The risk-side speaker granted that AI might "hit a wall," or that one of AI 2027's steps might go too far, and said: "I hope and pray that's true." The same speaker still described putting an agent swarm "10,000 strong" to work on a better AI architecture. The result would not need to be dramatically smarter: "It just has to be a little bit better at getting better." The speaker said they would bet against recursive self-improvement starting that soon. But given the swarms and the Millennium problems, they said, "it's kind of hard to have less than 1% in six months."