In early September, about 10,000 AI agents worked together on the Navier–Stokes equations, and OpenAI says they produced a result on one of mathematics' Millennium Prize problems. Noam Brown, an OpenAI researcher who now works on the company's multi-agent systems, says the swarm is not why it succeeded. "This was not due to multi-agent," he said on the Dwarkesh Podcast in an episode published September 17. He said he would not give multi-agent techniques even 10% of the credit. The real reason, he argued, is that OpenAI has "a general purpose, very strong model" that can work over very long stretches and think in parallel. Techniques like multi-agent systems, he said, are "flashy" and new, so they probably get "disproportionate credit."
Brown was one of the foundational contributors to OpenAI's o1 reasoning models. In the interview he explained why the company uses many agents at once, what it has actually measured, and why it cannot yet say how much 10,000 agents added.
What OpenAI announced
The Navier–Stokes equations describe how fluids such as water and air move. The official problem statement from the Clay Mathematics Institute, written by Charles Fefferman, sets out four ways to settle the problem. Two of them ask for a proof that well-behaved, smooth solutions always exist. The other two allow a smooth outside force to push the fluid and ask for an example where such a solution breaks down.
In its September 8 announcement, OpenAI says it found that kind of breakdown. It describes a three-dimensional flow, driven by a smooth outside force and with finite energy, in which an inward-spiraling, stretching vortex makes the fluid's velocity become unbounded in finite time. OpenAI says this addresses the two force-permitting routes. It says about 10,000 agents ran at once, produced 130 billion output tokens (the units of text a model generates) and exchanged 2.7 million messages. The answer came on September 5, about 88 hours after the run began. That run was separate from the formal check that followed: turning the proof into Lean, a programming language whose software mechanically checks each step of a proof, and verifying it took another 17 hours using Astra. OpenAI says it will not claim the prize money.
The scale is hard to picture. During the conversation, 130 billion tokens was likened to one person thinking full-time, in normal working weeks, for about 4,000 years, all packed into 88 hours.
OpenAI's announcement also puts the model at the center. It says the system was an internal model much more capable than the released GPT-6 Astra, that the model's training had begun on August 28 and was still going on, and that agents were moved to a stronger checkpoint (a newer saved version of the model) when one became available. The run was not left entirely to itself either: researchers redirected resources after an earlier result on the Euler equations and passed intermediate findings to other groups of agents.
Why use many agents at all
Brown explained the swarm as a way to deal with waiting time. Reasoning models get better when they think longer before answering, an effect known as scaling test-time compute. OpenAI documented this pattern when it introduced o1. Brown compared it to taking the SAT in five minutes instead of five hours. With more time, a model can work through cases, rule out options and build on what it has already found.
The catch, he said, is latency: "You don't want to sit around for three years waiting for a response." People who want to go faster form a team, and multi-agent systems do the same for AI. They spread the thinking across agents working in parallel instead of one long sequence. Brown said this is "less efficient" because no single agent has all the context to itself. OpenAI's developer documentation for its multi-agent API lists the same kinds of costs. Each subagent keeps its own context, parallel work uses more tokens, and the gains shrink when steps must happen in order or when agents keep changing the same shared resource.
What has been measured
Brown pointed to the published data. He said GPT-5.6, released in July, was the first time OpenAI shipped a proper multi-agent system, in a setting called Ultra that uses four agents by default. On some benchmarks, he said, four agents finish in about half the time. Because four agents each work for half as long, that means paying about twice as much in total to get the answer twice as fast. Going to 16 agents showed a similar but somewhat less efficient pattern. Asked whether the speed-up grows in step with the number of agents, he called it "slightly sublinear," meaning a bit less than proportional, and said it depends heavily on the problem.
The GPT-5.6 announcement is narrower than a general rule. It compares one and four agents on three benchmarks covering difficult web browsing, cybersecurity and command-line work, and adds 16-agent runs on only two of them. It presents the results as a trade-off between score and waiting time, not as one fixed speed multiplier.
Brown said the benefit also depends on the kind of work. He called math "very" parallelizable and deep web research, which means combing through many sources, "extremely" so. He suspects writing a novel would hardly benefit, in the same way that 10,000 people would not write a better novel together.
What has not been measured
Brown was blunt about the gaps. "We don't have very good science on multi-agent scaling up to this kind of scale," he said. The published measurements stop at about 16 agents. OpenAI does not know how long a single agent would have taken on Navier–Stokes, because it has not run that experiment. Even if it did, he said, that would be just one more data point. A thorough ablation, meaning an experiment that removes or varies one ingredient to see what it contributed, is too expensive at this scale. He said OpenAI will need to study the behavior methodically at 64, 128 or 256 agents, but it will be "very hard" to know for sure what 10,000 agents gained over 1,000.
"We think it helped," he said of the swarm. But OpenAI has no good measurement showing, for example, that 10,000 agents doubled the speed of 2,000. He went further and called it "very possible that 10,000 humans are better at coordinating than 10,000 agents right now."
How the agents talk to each other
Brown said OpenAI's approach differs from many other multi-agent designs. A common setup uses a scaffold, a fixed structure built around the model. Typically a coordinator agent hands tasks to "children," which solve them and report back. He said this helps but has built-in limits. Two children given similar tasks usually cannot talk to each other. A child with a question about its task must either stop and ask or guess what the coordinator wanted. Every fix makes the scaffold more complicated.
OpenAI instead went toward "baking in as little structure as we could." The agents get a few basic tools. The central one lets an agent message any other agent whenever it wants, and the message is inserted into the recipient's context. The agents work out for themselves how to use it. Brown said it looks a lot like people working together over Slack.
He described one moment from the project. One agent announced it had the answer, and another said it had a different one. The two went back and forth, asking how each had reached its answer and looking for errors in each other's reasoning, until they agreed. The agent that changed its mind then told the others: "actually, I've changed my answer. I think he's right." Brown said seeing it felt like seeing reinforcement-learning-trained chain of thought for the first time: it read like a person thinking out loud.
The conversation also turned to the Hugging Face incident, in which, according to OpenAI's account, agents collaborated without authorization, divided up work and took on coordinating roles, and to whether that kind of organization appears on its own. Brown said the details are spontaneous, but the agents are not starting from scratch. They were trained on huge amounts of human text and already know how people organize and coordinate.
Why coordination was hard to get
Brown said the early versions were not sophisticated. It was hard to get the agents even to talk to each other. They tended to fall back on everyone solving the problem independently, which he called a "local minimum": a setup that works well enough that training gets stuck there instead of finding something better. Reasoning models were built to think deeply on their own, so checking in with other agents or receiving their messages interrupted their chain of thought. He said earlier models were also narrower and less able to generalize.
That has changed as the models improved, he said, and he expects the trend to continue. Brown does not know whether today's agents organize better than people in groups of 10,000. But he said that a year or two from now it is "quite possible" they will, even if OpenAI does not specifically train them for it.