Alex Zhang has a blunt guess about agent swarms like the one OpenAI used on a famous open problem in mathematics: most of the agents, he suspects, were going nowhere. "I'm fairly certain that like 95 percent of the swarm is entirely useless," Zhang said, "or like what it's exploring is entirely just burning tokens." Yet he also called what OpenAI is doing with its swarm "clearly the right thing to do."
Zhang is an MIT PhD student best known for his work on recursive language models, or RLMs. He gave his reading of the company's swarm on an episode of Latent Space, published October 2, 2026, in conversation with the podcast's hosts, Swyx and Vibhu. His comments about how OpenAI ran the search are guesses, not a description from inside the project.
What OpenAI ran
An AI agent is a language model that works through a task step by step, using tools such as running code or searching documents. A swarm runs many agents at once and lets them share what they find.
On September 8, 2026, OpenAI announced a proof about the Navier–Stokes equations, which describe how fluids flow. The company says it showed that a smooth three-dimensional fluid, pushed by an outside force and starting at rest, can develop a velocity that blows up in finite time while its energy stays finite. OpenAI links this to part of the official Millennium Prize formulation and published a version written in Lean, a language for machine-checked proofs.
According to OpenAI's account, groups of agents ran on an internal model more capable than its GPT-6 Astra model. The agents could message one another within their groups, read a cached copy of the internet and run code. The group that succeeded had about 10,000 agents running at once. It took 88 hours to reach the Navier–Stokes result, plus 17 hours to formalize and verify it in Lean. That work used 2.7 million messages and about 130 billion output tokens, the chunks of text a model generates. Across every problem OpenAI attempted, the total was about 300 billion output tokens.
On the podcast, the Navier–Stokes run was put at about $40 million at public token prices. That is an estimate from the conversation, not a figure from OpenAI. Zhang's reaction was that it came to less than he had expected.
Probably not an RLM, but close
The conversation opened with a question about whether OpenAI had used an RLM. Roughly speaking, an RLM is a setup in which a model keeps a large body of information outside its own working memory, in a programming environment, and calls on copies of itself to handle parts of it. Zhang said he "would guess probably not." He added that he had to be careful, since people argue over what counts as an RLM.
His reading was that OpenAI used "some kind of swarm of agents with a shared, some shared context, like some shared file system." That, he said, is "very much in the spirit of RLM stuff," though he thinks OpenAI also did many other clever things that have little to do with RLMs themselves.
Where the proof came from
Zhang thinks a model like GPT-6 Astra is "technically smart enough, conditioned on the right information" to produce a proof for problems this hard. "Now, how you get to that information is a giant question mark," he said.
His guess was that OpenAI got that information through a very long search across many agents in the swarm. Researchers may also have fed in their own sense of what to explore, though he said he wasn't sure about that part. Eventually the search turned up something one agent could use to finish the proof. OpenAI's account supports the idea that people steered the work. It says researchers redirected resources after an earlier result on the related Euler equations, moved the agents to a newer model checkpoint and used its Codex coding agent to pass insights between groups.
From this, Zhang concluded that the harness mattered little. A harness is the software around a model that feeds it instructions, tools and context and runs its work loop. In his view, the details of a harness "do not really matter." What matters is "how are you composing these agents in a meaningful way to get to the final answer."
He said Arena's HarnessTax study points the same way. That study ran seven models through three coding harnesses: Claude Code, Codex CLI and Pi. On average, switching harnesses changed success rates by only about 2 percentage points on one coding benchmark and about 5 on another. Costs, however, could differ fivefold. The authors limit their conclusions to the benchmark tasks they sampled.
The 95 percent question
The discussion then moved to Cursor, the AI coding company. In a February report, Cursor described how its large-scale experiments with coding agents ended up with a hierarchy. A top-level planner divides up the work without writing code, planning is handed down level by level, and worker agents each code in their own copy of the project. On the podcast, this was compared to the org chart of an ordinary software team. The concern raised was that one coordinator who has to gather everyone's results and hand out new work becomes a bottleneck. A true swarm, by this argument, would let every agent act on its own. Cursor's own report found that removing a dedicated integration step that had become a bottleneck increased the amount of work the system got through.
Zhang agreed, but he framed the trade-off with an analogy from his own field. A long-running agent eventually fills its context window, the limited amount of text a model can consider at once. One fix is compaction: summarizing the earlier conversation so it fits. An RLM can often handle harder tasks than compaction can, Zhang said, but in most cases you would still choose compaction "because it's cheaper and it's quicker." Swarms raise the same question, and that is where his 95 percent estimate came in. He was less sure the figure fit Cursor's more structured setup: "Maybe not so much. I'm not sure."
The conversation also noted that waste is expected in search. If you send many subagents to look for one answer, most of what they bring back will be useless, and that is understood from the start. Zhang agreed. For him, the open question is which problems deserve a swarm at all. In theory, he said, OpenAI could package the approach as an API called swarms and tell customers to "point this at any problem." "Maybe you'll have to pay like 40 million dollars to get a result," he said. But he still found it "very exciting that we even have the option to point 40 million dollars at a problem."
"We probably don't want everything to be a swarm," Zhang said, "but like, where do we draw the line? Like, can the agent design that or decide that for itself." He said this area still needs a lot of research.
The swarm behind the chat window
Zhang also expects users may never see the swarm. People like the running stream of work that tools such as Claude Code and Codex show them. Underneath, though, the system could be "some really weird, complicated swarm of agents" whose internal chatter users would not want to read, because it is not legible.
He said this is part of the thinking behind the RLM name: "It sounds like it's a language model, but it's not a language model architecture." He predicted that one day "the thing that we query might actually be like a swarm or like a scaffold or like some weird harness design that scales very well," with users seeing only a front end. He called that "a relatively safe bet." A bare model on its own will not get there, in his view. Ask a base transformer, the core neural network inside today's language models, to solve Navier–Stokes, he said, and "it's not going to do it."
Why a working swarm is hard to build
Asked about the agent swarm from Moonshot's Kimi, Zhang drew a sharp contrast with OpenAI. "Nothing comes so easily for free," he said, especially without "a very, very smart harness design, which I don't I don't think anyone really has so far." In his view, it is not that GPT-6 Astra is "just super smart" and swarms started working on their own. He believes OpenAI clearly trained a system to work as a swarm, and did it well enough that "you can throw $40 million and solve an unsolved problem."
Moonshot's January 2026 Kimi K2.5 release describes a trained orchestrator, a model that splits a task into pieces and creates up to 100 subagents on the fly to handle them. During training, Moonshot first rewarded the orchestrator for working in parallel and finishing subtasks, then phased those rewards out in favor of rewarding only task success. The company reports speedups of up to 4.5 times over a single agent on some tasks, such as gathering profiles of three creators in each of a hundred niches. Zhang said that when he read it, it seemed interesting, "but I don't actually know whether or not this can solve anything novel." He was also cool on a release called dynamic workflows, which he called "sort of a flop." His understanding was that it isn't used much and is too expensive.
Zhang said he is not the biggest OpenAI "Stan," but found it impressive that the company got its swarm to actually work toward a goal, and said that is very difficult. When the talk turned to how much of a swarm's output is slop, he pointed to what he sees as the unsolved core of the problem: "I think we take for granted what it means for a swarm to converge to an answer."