29 September 2026
Heard in AI

Schulman expects bigger AI models; O’Neill expects a plateau

John Schulman expects frontier AI models to keep growing as compute and GPUs expand. Charlie O’Neill expects little parameter growth for a few years, because long reinforcement-learning runs reward cheap inference.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Dwarkesh Podcast, episode published 11 September 2026

Charlie O’Neill does not expect the largest AI models to grow much over the next few years. John Schulman expects them to keep growing. The two researchers disagreed during a panel on the Dwarkesh Podcast. The episode was published on September 11, 2026, and Beren Millidge and host Dwarkesh Patel also took part. The disagreement was less about ambition than about which resource runs short first: computing power for running models, training tasks hard enough to need a bigger model, or the high-quality text that models learn from.

Two ways to count a model's size

A language model's size is usually given in parameters, the adjustable numbers set during training. More parameters give a model more capacity, but they also make it more expensive to run. Many frontier models are now "sparse". They are split into many specialist sub-networks, called experts, and each token (a word or piece of a word) goes to only a few of them. That creates two counts. Total parameters are everything the model stores. Active parameters are the ones actually used for each token, and they set most of the computing cost of each step.

The question put to the panel was about active parameters. It was suggested that frontier open models had grown roughly twofold a year, and that even closed frontier models might have something like 100 billion or 200 billion active parameters. Both figures were offered as rough estimates, not reported sizes. Would that doubling continue now that labs rely on reinforcement learning (RL), where they also want to save compute? Could models reach a threshold beyond which extra parameters matter less? And how many active parameters would a frontier model have in 2030?

O’Neill: long rollouts make size expensive

In reinforcement learning, a model attempts a task, gets a score for the attempt and learns from that score. Each attempt is called a rollout. O’Neill said labs are now focused on "longer and longer horizon rollouts": attempts in which a model works through a long task and generates a great deal of text along the way. Every token of that text has to be computed, so inference efficiency, meaning the cost of running the model, matters a lot.

In his view, models have not yet exhausted what they can learn from these tasks. The bottleneck is still the environments, the tasks and graders that labs build for training, so "we might see like a little bit of plateau" in size. He said he had a feeling that closed frontier models such as Mythos and the GPT models are much smaller than the 10-trillion-parameter range people talk about, and that comparing them with open-source models would probably support that conclusion. This is his inference. The sizes of those closed models have not been reported.

O’Neill described model size as a balance between several inputs. The first is how much pretraining data a lab has. The second is how difficult its RL environments are. Ideally, he said, a lab picks the size at which the model gets a decent pass rate on its first attempt at the hardest environments it has. Going beyond that means "paying like much more inference flops" (floating-point operations, a measure of computation) without any need. His forecast therefore depends on how quickly labs can make their training environments harder.

Schulman: compute grows, but data changes the calculation

Schulman took the other side. "I would expect the models to keep getting bigger just because people are scaling up compute and the GPUs are getting bigger," he said. How much bigger, he added, depends on scaling laws "in non-obvious ways". Scaling laws are the fitted relationships between a model's resources and its performance.

His main point was about data. Labs are reaching the point of running low on high-quality pretraining data. He therefore expects data efficiency, meaning how much a model learns from each example, to matter more than compute efficiency when labs choose architectures. That could change how sparse they make their models. "Parameters are a different resource than active parameters," he said. Sparsity has increased somewhat, but in his view it is not clear that it will keep increasing without limit: "There might be some kind of sweet spot." There is an argument that sparsity hurts data efficiency, because several experts may each have to learn the same thing, though he called that debatable. He does not think there is yet a good enough theory of scaling laws to explain why sparsity helps, how much it will help, or where its benefit levels off.

Asked to spell this out, Schulman described how architectures are usually compared. Each design choice produces its own scaling law. Traditionally, researchers plot performance against compute and keep the designs on the outer edge of that chart. If data goes on the horizontal axis instead of compute, because compute is plentiful and data is not, a different set of designs ends up on the best frontier.

Weights, caches and hardware

The discussion then separated the two parameter counts by what limits each one. The argument was that the need for cheap RL rollouts will push active parameters down considerably, while total parameters depend largely on hardware. Serving multi-trillion-parameter models requires very high memory bandwidth and a lot of GPU memory. Many labs still run Nvidia's H100 chips, and newer generations would make larger models easier to serve and larger RL workloads possible.

The distinction matters because a sparse model's idle experts still have to be stored in fast memory, even though each token uses only a few of them. Model weights are also not the only thing competing for that memory. While a model generates text, it keeps a key-value cache, which stores intermediate results for every earlier token so they need not be recomputed. The cache grows with the length of the context. The 2023 grouped-query attention paper lets several query heads share each set of cached keys and values. This shrinks the cache and the memory traffic of reading it, but the cache still grows with sequence length.

The trillion-parameter headline

O’Neill also doubted that models had really doubled in size every year. People have been training trillion-parameter models for years, he said. He cited a recent social-media post about an early experiment with a very sparse trillion-parameter model, which he attributed to OpenAI. Schulman corrected the history: the work had been done at Google, before its researchers went to OpenAI, and it was the Switch Transformer. O’Neill said that model was very good at knowledge but poor at reasoning because it was so sparse. By his account, the field has spent some time between 100 billion and 2 trillion parameters, and growth "certainly hasn't been" a steady line.

The Switch Transformer paper, published in January 2021, supports his broad point but qualifies his description of the model. Its Switch-C model had about 1.57 trillion parameters spread over 2,048 experts, with each token routed to just one expert. The sparse models trained efficiently and did well on knowledge-heavy question answering. Downstream results were mixed rather than uniformly poor. A different configuration, Switch-XXL, had fewer parameters but more computation per token. It trailed Google's T5-XXL model on the SuperGLUE language-understanding benchmark but improved on the ANLI reasoning test. A headline parameter count therefore says little about how much computation each token gets or how well the model performs.

What counts as efficiency

The size debate followed an argument about where recent progress has come from. Patel described an investigation with Princeton student Jerry Han. They trained model-training recipes from 2019 to the present on training datasets from the same period, in every combination. The aim was to measure how much less compute each change needs to reach a given capability. On the panel, Patel put the gain from data at about 9x and the gain from architecture at about 3x, "at a very small scale". The published write-up gives 12.0-fold for data and 3.7-fold for recipes at its largest compute budget. Its recipes include optimizer and training changes as well as architecture. All runs used a 2,048-token context, and the authors note that the experiment does not capture long-context gains or inference efficiency.

O’Neill multiplied the episode's figures into a combined gain of roughly 27-fold. He compared that with an estimate he recalled from Epoch "or someone" of about threefold improvement a year since 2019, which would compound to more than 2,000-fold. The missing factor, he suggested, might indicate how much progress comes from post-training. The closest published estimate is a 2024 study led by Anson Ho covering 2012 to 2023. It found that the compute needed for a fixed level of pretraining performance halved about every 8.4 months, or roughly 2.7-fold a year, with wide uncertainty (4.5 to 14.3 months). That study measured pretraining quality alone, and its estimate includes data improvements. Patel offered another explanation: many efficiency gains depend on scale, and his experiment ran at extremely small scale. He said the team did not have enough compute to test whether data gains or recipe gains depend more on scale.

The discussion also questioned whether these gains simply multiply. The argument was that an architecture change such as grouped-query attention lets a model reach a qualitatively new regime. Without such changes, handling a million tokens of context with full attention would be ridiculously expensive, so data at that length could never be used. A test with a 2,000-token context, where the architecture unlocks nothing, makes data look more important than it is. Much of today's mid-training and post-training data also becomes more useful with scale: a 100-million-parameter model trained on long agentic task traces would not improve the way a sensibly sized model does. O’Neill added that teams such as those behind Kimi and DeepSeek now design architectures around how the models will be used in the real world. The compressed attention in DeepSeek's models, for example, targets inference efficiency rather than lower training loss.

Straight lines that hide bugs

Schulman said it took people so long to find scaling laws in the first place because the relationship is clean only if everything is done right. The "beautiful straight lines on graphs" hide how carefully every hyperparameter has to be scaled with model size. O’Neill added that "bugs have their own clean scaling laws as well". He pointed to the earlier Kaplan scaling laws, which he said were thrown off by a learning-rate schedule and, he thought, by not counting embedding parameters, which make up a sizable share of small models.

A 2024 study by Porian and colleagues supports his broader point but qualifies the details. It traced the gap between earlier estimates and Chinchilla to three factors: how the final output layer's computation was counted, how long the warmup period ran, and how optimizer settings were tuned across scales. Correcting these moved the estimates toward Chinchilla. Careful learning-rate decay improved results, but it was not needed to recover the Chinchilla-like scaling relationship. Those experiments stayed below a billion parameters.

Chinchilla and running out of data

The Chinchilla study came up again when the discussion returned to model size. The 2022 paper trained more than 400 models and concluded that, for a fixed training budget, parameters and training tokens should grow roughly together. Its 70-billion-parameter Chinchilla model, trained on 1.4 trillion tokens, beat the 280-billion-parameter Gopher on most evaluations with a similar compute budget, and it cost less to run afterwards. The paper's target is pretraining performance for a fixed training budget. When heavy use after deployment is counted, it can pay to train smaller models for longer.

The argument on the panel was that labs are now on the "way too much data" side of Chinchilla: they deliberately over-train small models that are cheap to run. Larger models are more sample-efficient, meaning they generalize better and reach a lower error from the same data. If data becomes the constraint instead of compute, labs could move back toward the Chinchilla point, or even toward large models that have not absorbed everything they could. A counterpoint came in the same exchange: under the basic Chinchilla law, adding parameters reduces the data needed to reach the same loss only a little, by less than tenfold even with unlimited parameters.

O’Neill questioned the premise. Even with new chips arriving, he argued, labs will be so short of compute over the next few years that the shift won't necessarily happen. The reply was that it depends on the balance between compute spent on training and compute spent on inference. A lab that is badly short of data but not of compute should build bigger models. A lab short of compute should always build smaller ones. And compute can also be used to generate synthetic data, which makes the outcome, in the words of the reply, one of those "very hard things to predict".

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06

Connected ideas and articles

From the conversation

Podcast episodes

Dwarkesh Podcast

AI researchers debate how close we are to recursive self-improvement

Episode published This article draws on 1:06:13–1:18:01 (approximate times)

Article history

Updates to this article

Tags