When recursive language models (RLMs) first appeared, a common reaction was a shrug. What is even the purpose of this? Isn't it just subagents? Alex Zhang, an MIT PhD student known for the RLM work, recalled those responses on the Latent Space podcast as "almost like a good sign." If people can't see why something exists, he argued, they probably haven't thought about the problem it addresses.
That reaction is central to a broader argument Zhang made about how researchers in universities should choose their work. Academics, he said, should take big bets on problems industry labs are not looking at, including ideas that seem trivial, odd or pointless when they first appear. This is a view about research strategy, built from examples. It is not a measured finding.
What the argument says
Zhang said the most successful research from graduate students and academics tends to come from people who care about problems "that maybe like most people in industry are not looking at." He sees the opposite habit as common. Many grad students, he said, work on things that "look good to an industry lab," such as a currently popular benchmark or a harness built for one specific task. (A harness is the software loop around an AI model that lets it use tools and carry out multi-step work.) The reason someone would do that, he said, is a clear goal shaped around the models that exist today.
His reasoning is about comparative advantage. Industry labs have "tons of resources, tons of talent." Academia, he said, cannot afford to produce flagship releases. He contrasted the papers he admires with a release like OpenAI's GPT-6 Astra, which everyone rushes to use. If a PhD student works on the same problems as the labs, he asked, why accept fewer resources and fewer colleagues? "You kind of need to take big bets if you're going to be in academia," Zhang said. "Otherwise, I think like just go to an industry lab."
He also named what academics do have. A PhD student, he said, "can work on literally whatever you want for the most part," and doesn't "have to deal with bureaucracy and all these other things." He called that the student's biggest advantage over anyone at another lab. In the discussion, this was framed as an unfair advantage in a landscape otherwise biased against academics.
Where the view comes from
Zhang does not present research taste as purely innate. He said it gets "developed through opportunities," and that standout careers mix luck with being smart. He credits his own path. Before his PhD he worked at Princeton with the team behind SWE-bench, and at MIT he works with his advisor, Omar, whom he called "a fantastic advisor." He added that the same lucky streaks happen to people inside industry labs. Grad students, he said, are just more visible.
The examples he points to
SWE-bench. Zhang retold a story he said Ophir, from the Princeton team behind the benchmark, "loves to tell." When SWE-bench came out, "nobody cared." People thought the task was impossible and asked why it would ever count as a benchmark. Only after the coding agent Devin appeared, Zhang said, did everyone decide it was "something we want to hill climb." The original SWE-bench paper helps explain the skepticism. It asked models to fix 2,294 real GitHub issues from 12 Python repositories, and an issue counted as solved only if the model's patch passed the project's tests. In the paper's evaluation with a BM25 retriever, Claude 2 resolved 1.96% of issues.
Quiet-STaR, chain of thought and ReAct. Zhang called Eric Zelikman's STaR and Quiet-STaR work his favorite example. On first reading, he said, he wondered whether it was simply an obvious idea. He felt the same about chain-of-thought prompting and ReAct. Quiet-STaR trains a model to generate internal "thoughts" while reading ordinary text, and rewards the thoughts that help it predict what comes next. The paper is itself a modest bet. Experiments used a single 7-billion-parameter model (Mistral 7B). Direct-answer gains on math and commonsense tests came without task-specific fine-tuning, but gains were smaller on general web text. Generating the thoughts added substantial computing cost. ReAct, from Shunyu Yao and colleagues, alternates a model's reasoning with actions such as a Wikipedia search, letting each result shape the next step.
What makes such a paper valuable, Zhang said, is that "it tells a bit of a story as to like what you want the field to look like." The idea may look simple, but it points somewhere.
Recursive language models. Zhang called the RLM paper "a super, super simple idea." As the episode describes it, an RLM offloads context and hands work to code and recursive subagents. The early "just subagents" dismissals are what he cites as the encouraging signal.
Why he moved from GPU kernels to harnesses
Zhang's own path illustrates the bets he describes. He was deeply involved in GPU Mode, a community for writing GPU kernels (the low-level code that runs AI math on graphics chips), and in KernelBench. The discussion turned to whether harness work feels less legitimate than messing with GPUs. Zhang said it can be "very uncomfortable" for someone who likes to think in a math-oriented way, and that findings are hard to verify with the compute academics have. But he thinks harnesses are where "most of like the innovation is yet to happen." Kernels, for him, were a means: automating them serves a broader goal of exploring ideas "where I'm not bottlenecked by systems challenges." He said he is also interested in work at the model level.
The same logic shapes what he leaves alone. Zhang said he is not training RLMs at scale at MIT because he can't afford to, though companies are working on it. He wondered whether training models around a smart harness might reveal better post-training scaling laws. He said he would rather spend his PhD on other big bets.
How he filters ideas
Zhang said he has "10 or 15 different ideas," and that "most of them are bad." His method is a short look followed by full commitment. He spends some time on an idea, including ones people bring to him. If there seems to be something there, he takes "the next few weeks and just really pursue it." Much of the thinking happens away from the desk, on a run or playing tennis, which he called "the most fun times." Once convinced, he said, he will "drop everything and just do it," until experiments are running and it is "easy coasting again."
Objections and open questions
Zhang's argument has limits built in. Big bets usually fail. "A lot of them will fail," he said, and of the bets open to PhD students, "most of them will probably yield nothing." The harness work he chose is hard to verify on academic budgets. He said academia can't afford flagship releases "at least right now" and that there are "a whole slew of reasons" that should change, without listing them.
The discussion also tied the argument to career incentives, contrasting staying in school to pursue open-endedness with joining a lab to maximize profit. Zhang said he doesn't hear this discourse much because he is on the East Coast, but that whenever he comes "here," it is always the topic of discussion. He also said most progress in the field has been "a little bit boring." The outcomes aren't boring, he said, but the process often is.