October 2, 2026
Heard in AI

Inside Prime Agent, an RLM harness where Python is the only tool

Alex Zhang, an MIT PhD student, described Prime Agent, built with Prime Intellect: the agent acts only through Python and keeps its history on disk. He also cited outside RLM work, including Harvey's legal agents and Laude Institute's Headlong.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published October 2, 2026

Most AI coding agents give a model a menu of tools: one to read files, one to search, one to run commands. Prime Agent strips that menu down to a single item. According to Alex Zhang, an MIT PhD student best known for his work on recursive language models (RLMs), the agent's only tool is an IPython session, an interactive Python environment. Everything else the agent can do has to be written as code inside it.

Zhang described the design on an episode of the Latent Space podcast published October 2, 2026. He also explained how he came to work on it with the company Prime Intellect, and pointed to outside groups that have taken his RLM idea in their own directions.

What an RLM is

A language model on its own only reads text and produces text. To do longer tasks it runs inside a harness: the surrounding software that hands it a prompt, runs the tools it asks for, feeds back the results and keeps track of the conversation. Products such as Claude Code and Codex are harnesses of this kind.

Zhang offered a compact definition of an RLM, saying he wished he had put it in the original paper. An RLM, he said, "is basically just a harness design where the only tool in the harness is code." Other tools exist, but as functions the model calls from code. One of those functions is the model itself, so it can hand pieces of a job to copies of itself, called subagents, in what Zhang called "programmatic subagent calling." And the material the model is working on is always stored somewhere in that code environment, such as a file system, rather than being stuffed entirely into the prompt. Moving material out of the prompt in this way is known as context offloading.

How Zhang got involved

Zhang said Prime Intellect, which he stressed was not affiliated with him, published a blog post arguing that RLMs were the future. The likely post is Prime Intellect's January 2026 article calling RLMs "the paradigm of 2026." Its own early tests were mixed: RLMs helped with long documents and tools that return large outputs, but could do worse on math or verbatim copying when models did not use the setup well.

A friend from GPU Mode, the online community for GPU programmers, Matei, worked at the company and put them in touch. What impressed Zhang, he said, was that Prime's researchers understood the point of the RLM paper. It was "not necessarily just to say that like, you know, we're solving long context tasks," he said, "but actually like we want more opinionated harness designs."

After deciding to work together, they set out both to train an RLM and to build an RLM harness, and that, Zhang said, "is how like Prime Agent came about." His main worry at the start was that none of the available models were very good at working this way. In the end, he said, they "got kind of lucky": many new frontier models work well inside it, as do some of the newer open-source ones.

The design

Zhang said Prime Agent is built on top of Pi, a minimalist harness he uses as his reference point because, in his view, "all other harnesses are basically just Pi." The difference is that IPython is the only tool. Every other tool is loaded as a Python module or a bash script the model can run.

The second key feature is how Prime Agent handles memory. Long agent sessions eventually exceed what a model can hold in its prompt, so harnesses compact them: they replace older parts of the conversation with a summary. Zhang said Prime Agent still uses that standard compaction loop, but stores the full trajectory, the complete record of what the agent has seen and done, on disk. "So the model can always reference its original context, even if it's compacted," he said. He described this as context offloading, "but not all the way." Prime Intellect's announcement of Prime Agent describes the same arrangement: the entire session history is stored as append-only files on disk, and compaction is mainly used to clean the agent's main context, while the full history, including past compactions, can be accessed programmatically in the IPython kernel when needed.

Third is what Zhang called a continual harness, an idea from Seth, another PhD student, who used it to get language models to play games. It is a rule about which parts of the harness the harness itself may change. Zhang said it can rewrite its own skills (reusable instruction packages), the subagents available to it and the system prompt. It is exposed as one more tool inside the Python session. The repository is more cautious about the system prompt: refinements based on evidence from past runs add notes, memories, skills or subagent specifications, while the base system prompt stays fixed. Changes can be rolled back.

Fourth, because RLMs tend to spawn many subagents, Prime Agent has a framework for subagents to message the main agent or each other. There are rules about which subagent may talk to which, and, true to the design, the model writes the code that sends the messages. Finally, subagents can be persistent. They can outlast the run of the agent that created them, and a user can go back into one and prompt it further, which Zhang said gives "more visibility and flexibility" into what is going on.

That feature drew an immediate response in the conversation: Codex's subagents were described as "very ephemeral," with users discouraged from giving them long-running work. Zhang said that made sense. The discussion settled on externalizing state to a file system as the way around it, which Zhang called "the big trick," along with forcing everything through code and trusting a model that can write code to build its own harness.

The project's repository adds a caution for anyone running it: the agent's processes are isolated from one another, but that is not a security sandbox. Code the agent executes otherwise runs with the user's own permissions.

The announcement also reports benchmark tests run without a model trained specifically for the harness. Across long-context benchmarks, its gains depended on the task, and some combinations of model and harness did as well or better outside Prime Agent.

Training, and where Zhang steps back

Zhang said Prime Intellect is training a model of its own for this setup, something he said the company was public about in March, and is showing it can do so on its hosted training stack. He is not involved in that part. "I have other things in the PhD I want to work on," he said, adding that one of the luxuries of being a PhD student is the number of big bets available, "most of them will probably yield nothing."

RLMs in other hands

Asked which outside work deserved attention, Zhang began with Harvey, the legal AI company, which he said had post-trained an RLM on legal work without telling him in advance. He said that kind of work often means sifting through many documents for details that a pure retrieval system would struggle to find, and that the results were "really, really good."

Harvey's report, written with the AI infrastructure company Baseten, supports that, with conditions. Its agents wrote cited due-diligence memos for mergers and acquisitions using synthetic data rooms, collections of up to 5,000 documents and 80 million tokens, and the memos were judged against hundreds of criteria set by experts. Across seven models on 50 held-out rooms, the average share of criteria met rose from 23.3% with a standard tool loop to 62.4% with RLMs. Training also helped. One open model trained with reinforcement learning rose from 29.9% to 63.0% on held-out rooms. The report also found costs: six of the seven models were more expensive to run as RLMs. Letting subagents spawn their own subagents made results worse in ten of fourteen rooms tested.

Zhang's second pick was Headlong, from the Laude Institute, which he called Andy Konwinski's big project. He said it uses the RLM abstraction but does "something much cooler": a system that "thinks persistently," working through problems in its context even when no one is asking it anything. It is always on, he said, but controls token spending so it does not burn through a user's credits. Headlong's repository describes a thought loop that keeps running between conversations, slows down when no messages arrive and speeds up when they do. Its maintainers report costs of $1 to $2 an hour with their own settings, depending on the model and loop speed. The repository labels the project alpha research software and notes that multiple users share one timeline without strict privacy separation.

Zhang also mentioned Axe, which he said is a one-person harness built around the programming framework DSPy and RLMs, and noted that DSPy has an RLM of its own. His last example came from ARC-AGI-3, a benchmark of puzzle-like games. He said many harnesses in the benchmark's official Kaggle competition claim to use, or cite, some form of the RLM abstraction. That, he said, is exactly the kind of setting where combining code and symbolic reasoning with AI should pay off.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

From the conversation

Podcast episodes

Latent Space

Academia is for Ambition — Alex Zhang, MIT

Episode published This article draws on 44:24–45:58 and 52:01–1:00:30 (approximate times)

Article history

Updates to this article

Tags