29 September 2026
Heard in AI

Researchers explain how pass/fail rewards can yield big AI gains

On the Dwarkesh Podcast, AI researchers argued that pass/fail rewards work because earlier training supplies candidate behaviors and rewards ignore incidental wording. Charlie O'Neill warned that harder training tasks cost ever more to build.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Dwarkesh Podcast, episode published 11 September 2026

A pass/fail reward tells an AI model very little. After a long attempt at a maths problem or a coding task, the model learns roughly one thing: whether it got the answer right. Yet models trained this way have become much more capable. On an episode of the Dwarkesh Podcast published on September 11, 2026, host Dwarkesh Patel asked three AI researchers to explain that contradiction. The guests were John Schulman, Beren Millidge and Charlie O'Neill.

Their answers begin before the reward arrives. Earlier training already gives the model behaviors it can try. The reward points to what worked, without forcing the model to copy every word of a correct solution. The training tasks are arranged like a ladder, so that success stays within reach. O'Neill added that this ladder becomes more expensive to climb with each rung.

The puzzle: learning from almost nothing

The method in question is reinforcement learning (RL). A model attempts a task, an automatic checker scores the result, and training makes the successful behavior more likely. Patel recalled that a year earlier, many people had argued RL would not scale well. He noted that Schulman had written that a model learns about one bit per episode: did it get the answer right or not?

In Schulman's September 2025 study LoRA Without Regret, published by Thinking Machines, the arithmetic is concrete. About 10,000 maths problems with 32 attempts each produce roughly 320,000 right-or-wrong outcomes. That is far fewer than the roughly three million adjustable numbers in even the smallest adapter the study tested. (An adapter is a small add-on that is trained while the main model stays frozen.)

Patel's own essay from November 2025 argued that the problem is worse than it looks. A yes-or-no outcome carries a full bit only when success and failure are equally likely. When a model almost never succeeds, an attempt teaches it almost nothing, even though long attempts are expensive to run. "But I look at the models today, and they seem pretty smart," Patel said, "and it seems to be the result of scaling up RL."

Answer one: the behaviors were already there

The first answer pointed to a stage the discussion said is often underestimated: mid-training. This stage resembles pretraining, the initial phase in which a model learns from huge amounts of text. In mid-training, however, the material is synthetic reasoning data and environment-like tasks that "warm start" the model for RL. The discussion estimated that mid-training often takes a model about 80% of the way to its final RL checkpoint. On that view, RL mostly adjusts the model's policy, its tendencies to choose one behavior over another. It does not have to learn everything from scratch, so a few bits per episode can be enough.

Millidge makes the same argument in a July 2026 essay. He proposes that pretraining and mid-training supply much of the knowledge and many candidate behaviors, and RL concentrates probability on the useful ones. He presents this as an interpretation of how the process works, not as the result of a controlled experiment.

Answer two: the reward ignores the noise

The second answer compared RL with supervised fine-tuning, in which a model is trained to reproduce example solutions token by token. A worked maths solution used for fine-tuning also ends with the answer, so the crucial bit is present there too. The difference, as the discussion put it, is that fine-tuning also makes the model match the exact reasoning tokens of whatever produced the example. That gives it too many bits about one particular way of reasoning. RL's objective ignores all of that and rewards only the outcome. The useful signal is therefore not drowned out. The discussion called this a dramatic increase in signal-to-noise ratio, which is why RL can be so efficient per training step. Millidge's essay draws the same distinction: a task reward can target the behavior that matters, while next-token imitation must also model incidental wording.

The conversation then pressed further. If RL mainly upweights ways of thinking the model already had, how does that square with how much more capable models seem? The reply was that a small change to a model's parameters need not have a small effect on what it does. Even a single bit can sharply change how the model maps inputs to outputs, for example by ruling out half of the possible hypotheses. Schulman's LoRA study fits that picture. In its maths RL experiments with an 8-billion-parameter Llama model, even the smallest (rank-one) adapters matched full fine-tuning of the model, within experimental noise.

Longer work, not broader reasoning

O'Neill added a distinction of his own. Many people hoped RL would teach reasoning that carries across fields, which he called horizontal generalization. He does not think that happened: "just training on math doesn't necessarily make you the greatest coder." Labs still have to train on coding environments. What models did gain, he argued, is horizon generalization. Trained on progressively longer tasks, they learned to spend more tokens and keep making progress. Placed in an unfamiliar environment, they may lack the right reasoning patterns but can still keep going for longer, and that persistence correlates with success. The discussion also noted that RL does transfer somewhat, between maths and code or between puzzles and maths, and that labs now train on a much wider range of everyday tasks than they did two years ago.

As evidence, O'Neill cited EdgeBench, which he said showed that the rate at which models can work for longer is doubling every three months. The EdgeBench paper measures something different. Its roughly three-month doubling concerns learning speed: how much the leading models improve within two hours on a fixed set of 18 tasks, compared across model release dates. It is not a doubling of how long a model can work autonomously.

O'Neill then explained why progress can feel like a qualitative leap. In pretraining, he said, a smooth overall learning curve hides many sudden, separate gains. In one example, a model suddenly develops "induction heads," internal circuits that help it continue patterns it has seen earlier in a text. He thinks RL behaves similarly. On a single finance or Excel task, a model can jump from a 0.5% to a 90% pass rate. Averaged over many tasks and combined with longer working horizons, these jumps look like much better models. He described a slower outer loop around all this, which he said Millidge had mentioned. Labs train a model and apply RL, then add its synthetic reasoning traces to the mid-training data of the next model. Each generation's RL successes become part of the next generation's starting point.

Domain by domain

That helps explain an earlier part of the conversation about why labs train models on specific kinds of work. O'Neill described Anthropic's training environments as a clear example. The company started with coding, probably the lowest-hanging fruit because of the data available. It then moved to finance, with "so much Excel data," then presentations, and on into the long tail of office work. He said other labs, including open-source ones, have concluded that this was the right bet.

Schulman agreed that providers are going "domain by domain" and strengthening the highest-value areas. He called this one answer to why models have improved so much. In theory, he said, a model that learned well enough from context could read the finance books on the fly. It would not need finance training at all. Even then, labs might still use RL to build these intuitions into the model's weights so that it runs more efficiently. The discussion added that large models have room to learn many domains at once. It also noted that there is little ready-made data on doing AI research itself, so borrowing transfer from other work is worthwhile.

A ladder whose rungs get harder to build

O'Neill thinks some ladder of RL environments could probably produce an AI researcher at least as good as a human one. But he said the effort to climb each rung grows "kind of exponentially." For now, environment builders use shortcuts where a task is easier to set than to solve. One shortcut is to hide a complicated process that generates data and make the model work out what it is. Another is to take a hard-won real discovery and repackage it. His example was a bug that Anthropic found through a large combined effort by humans and language models. Turned into a training environment, it is something a single model could, in principle, find within a few million tokens. Eventually, he argued, people will have to build long tasks by hand. The agents will also need time and computing power to work through them, and he expects returns to diminish.

The discussion explained why the rungs matter. RL is currently poor at exploration. If a model cannot succeed within about 128 attempts, it is unlikely to get any signal to improve. RL therefore needs curricula, sequences of gradually harder tasks, which pretraining does not. Patel's essay also treats curricula as one way to keep success rates in the range where rewards carry information.

The researchers used two experiments to illustrate the difference between copying and climbing. Schulman described Talkie, a model pretrained only on text up to and including 1930. After fine-tuning on a moderate amount of modern coding-agent data, he said, it beat Claude 3 Opus, a much larger model, on SWE-bench, a test of fixing real software issues. For him, this showed that once expert behavior exists, "it's actually surprisingly easy" to copy it into a relatively weak model. The project's repository reports a training run that used about 75,000 example agent trajectories. On a 446-issue subset of the test, the 13-billion-parameter pre-1931 model solved 4.48% of issues on its first attempt, averaged over five evaluation runs. A sibling model of the same size pretrained on web text, given the same fine-tuning, solved 5.75%.

O'Neill offered a counterexample. He described a paper in which researchers trained a model only up to fifth-grade maths and primary-school English. They then tried to use RL to teach it late high-school and college maths, and he said "the gap was just too large" for it to climb. Moving up year by year, he argued, would have worked. The likely paper is LittleLearner, whose models learned only from kindergarten to Grade 5 material. Its results are milder than his summary. The restricted models made modest gains beyond elementary maths, not none, while a control model pretrained without the restriction gained much more. The authors also found that school grade order does not consistently match what models find difficult. Their tests covered models of 0.6 to 5 billion parameters.

Where the next bits come from

O'Neill's larger worry is supply. The world produces plenty of signal, from spreadsheet work to legal tasks. But he asked how much of it lies beyond what current models can already do. How many new maths or coding problems are being solved that current models could not solve? Even the world as a whole, he said, may not provide the bits needed to tip models into the next "basin of capability."

The discussion drew the contrast with pretraining. There, the needed signal already sits in large public collections of web text such as Common Crawl, and the work is mostly filtering out noise, which can be largely automated. For later training stages, the signal does not exist in the original data: no amount of filtering will uncover a hidden proof of a Millennium Prize problem. It has to come from elsewhere. Humans can write out their reasoning, humans can decide what environments to build and what they should reward, or models can learn from human data generated when they are deployed.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06

Connected ideas and articles

From the conversation

Podcast episodes

Dwarkesh Podcast

AI researchers debate how close we are to recursive self-improvement

Episode published This article draws on 33:50–38:47, 1:00:10–1:06:13 and 1:18:02–1:24:54 (approximate times)

Article history

Updates to this article

Tags