Developers argue about whether Claude Code, Codex or Pi is the best tool for letting an AI model work through a software project. Alex Zhang, a PhD student at MIT and an author of the research on recursive language models (RLMs), thinks the argument misses the point. "To be honest, I think all of them are the same," he said on Latent Space. "Most of the design decisions or like the design choices around these harnesses are the same."
His broader claim is that the field has spent too little effort designing the software that wraps a model, and that a different kind of wrapper could help models carry what they learn on one task over to tasks that look nothing like it.
What a harness does
A language model on its own predicts the next piece of text. Zhang said that format is "a really, really awkward form" for long, difficult jobs. His example was SWE-bench, a benchmark built from real software bug reports: you cannot solve a query over a whole code base with a single model call. A harness is the program around the model that makes this possible. It runs the model in a loop, gives it tools such as file search and command execution, and feeds the results back in.
Zhang described a harness as "a very, very opinionated program over how you want a language model to be form fit over a problem." The useful question, he said, is which of the harness's choices actually let the model solve the task.
By that standard, he sees little difference among the popular coding agents. When the logic is broken down, he said, Claude Code, Codex and Pi are "virtually the same thing." All of them use a pattern he calls "trajectory as a prompt." The full record of the model's actions and results so far is kept as the context the main model reads at every step, and it is compressed when it gets too long. He said this remains true even when such harnesses spin off helper agents, or subagents. He called Prime Agent "a little bit different" because it is built as an RLM, and said Grokbot also seemed quite different from the rest.
A study that fits his prior
Zhang pointed to a recent study by Arena, which he said "puts forward a prior that I had": most harness choices don't matter. Arena's HarnessTax study, published September 16, 2026, ran seven models through Claude Code, Codex CLI and Pi on samples of 30 tasks each from two coding benchmarks, SWE-bench Lite and Terminal-Bench 2.0. Average success differences linked to the harness were small: within about two percentage points on SWE-bench Lite and about five on Terminal-Bench. Costs, by contrast, could differ fivefold. That fits Zhang's remark that the main difference between these harnesses is "just cost for the most part." The study's authors limit their conclusions to these sampled open benchmarks.
The conversation raised the obvious objection: simple, general harnesses that are good at code are already being used for much more than software, and open models need to be trained inside a harness to use it well. Zhang agreed that training matters, and said he was "pretty sure" Anthropic and OpenAI train only on their own harnesses. He framed that as an assumption: "I'd assume not because I don't know why they would do that," he said of training on competitors' tools. As models get stronger, he argued, the difference shrinks. Putting OpenAI's Astra model inside the open-source OpenCode harness, for example, would not make it "go crazy."
That training habit cuts the other way for new designs, in Zhang's view. Frontier models placed inside an RLM are only "OK," he said, because they have been trained around the trajectory-as-prompt loop. He said a model called Fable was for a long time the best model for RLMs because it had been trained on dynamic workflows. It worked well even when it was "still like a little dumb." He said Astra is now good enough too. Asked whether that came from published results, he said they were internal and not public.
What an RLM is
Zhang's own definition, which he said he wished had been in the original paper: an RLM is "a harness design where the only tool in the harness is code." The model writes and runs code, for example in a Python environment. Inside that code it can call other tools, including fresh copies of itself, as ordinary functions. Zhang calls this programmatic subagent calling. The material the model is working on is kept not in its prompt but in memory inside that code environment, such as variables or files on disk. Even after its working history has been compressed, the model can always go back to the original.
Zhang's original RLM write-up shows how this works with a long document. Instead of reading the whole thing, the model can search it with code, split it into chunks, send each chunk to a subcall, and keep the answers in variables. For example, subcalls can label each entry, and code then counts the labels. That way, the size of the dataset is separated from how much any single model call has to read.
The idea was first presented as a fix for long inputs, which Zhang said harnesses handled badly. The December 2025 paper by Zhang, Kraska and Khattab reported large gains on some long-context tasks. It also found cases where simpler setups did better, such as smaller inputs. Zhang said his interest has since moved to composition. Ordinary tool calls must be invoked one turn at a time, with "no central context that you can kind of draw back from." An RLM gives the model a shared workspace and lets it write the program that ties steps together. He compared that workspace to the message boards that agent swarms use to communicate, and said the RLM's bet is that code is the best medium, "because these models are so good at writing code."
Why the same program can solve different tasks
The claim that RLMs help models generalize comes from a July 2026 post by Zhang and Khattab titled "Language model harnesses are compositional generalizers." Zhang walked through one of its examples: a retrieval task and an aggregation task in completely different domains. A regular model trained on both produces very different action records, or trajectories, for each. Inside an RLM, he said, the top-level strategy ends up looking the same. The subagents see different problems, but each handles an easier piece the model is smart enough to do. So training on one task lets the system "immediately solve" the other.
The same logic applies to length. Zhang said a strategy learned on short tasks transfers directly to longer ones because "they're effectively the same program," with only a length variable changed. The discussion cited tasks eight to 30 times longer than those trained on. The post's figures support that range: one benchmark goes from about 64,000 tokens of input in training to 2 million, another from 32,000 to 256,000.
The experiments behind those claims were modest. They trained a single open model, Qwen3-30B-A3B, over a few hundred training steps. Training as an RLM improved results on held-out lengths and domains, while gains from training the plain model transferred poorly. The authors also note that a model can defeat the generalization by shortcutting a task into one oversized subcall. They add that RLM training took roughly 1.5 to 3 times longer in wall-clock time. "There's no magic here," Zhang said.
He called the key property "locally in distribution." A task is "in distribution" when it resembles something the model has seen before. A long, unusual job is usually out of distribution as a whole. If the harness breaks the job into a program of small subagent calls, Zhang argued, each individual call can still look familiar. "If every task is in distribution for each individual language model call, you will probably get to the right answer," he said.
He argued that this works recursively. He believes competitive programming and GPU optimization (writing fast code for graphics chips) use similar skills. A trained RLM might learn the same plan for both: have subagents list promising solutions, then write a loop that tests and improves them against a checker. What the subagents do inside that plan may itself take the same form.
A primitive start, and a scaling question
Zhang does not present the RLM as the final design. He called it "a very primitive inductive bias," meaning a built-in assumption about how to approach problems. He said there is nothing special about it beyond being very different from what is used now. He expects "serious gains from very, very opinionated and good harness design." He also raised a stranger question: whether a model could be trained to act as an RLM "implicitly in its forward pass," inside its own computation.
Pressed on what a "smarter harness" could even mean, since he had just called the mainstream ones all the same, Zhang drew an analogy to scaling laws. These are the predictable curves relating a model's performance to its size and training data. Pretraining scaling laws hold, he argued, because today's architecture choices are fairly stable. Change the architecture completely and the power law would probably look very different. He thinks harnesses may be similar: today's mostly look alike, but radically different ones could change how post-training scales. The conversation added that changing how training data is represented could shift those curves too.
The discussion then pushed the idea to its logical end: train a custom model inside the harness, with the system collecting its own data like a fully automated AI researcher. Zhang agreed, but said that as an academic he is not doing scaled training of RLMs at MIT "because I can't afford to." He pointed to companies working on it, including Prime Intellect. Its January 2026 article reported mixed early results that varied by model and task, and proposed reinforcement learning as the next step.
"A skill issue"
The argument returns to a hypothesis Zhang published with Zhening Li and Khattab, the Mismanaged Geniuses essay. It proposes that frontier models are underused because the systems around them organize model calls poorly. The authors frame better management as a hypothesis about where the next capability jump could come from, not as a finding.
Put to him as "a skill issue," Zhang agreed. Current frontier models, even Astra, still cannot do one job consistently and well over the span of a month. He called that "a stupid problem" he believes is solvable without a frontier lab's resources. He thinks models are smart enough to match "some 18 year old high school kid doing some job," and that the language-model format is the obstacle, one a harness can be shaped around. Asked whether this was a point about capability overhang, meaning that current models could have much more impact even if progress paused, he said yes, "in some sense."
He left one experiment open. Remove all GPU programming data from a model as capable as Astra, he suggested, and see whether it could still learn enough in context to optimize GPU kernels inside the right harness. A human who knew as much as Astra could probably manage it, he said, and "I think we can actually approximate the human a lot more."