29 September 2026
Heard in AI

Schulman and O’Neill split on when AI speeds AI research tenfold

On a Dwarkesh Podcast panel, John Schulman forecast that AI will make AI researchers ten times more productive in about two years. Charlie O’Neill put it five to ten years away, arguing that choosing experiments and objectives is hard to automate.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Dwarkesh Podcast, episode published 11 September 2026

John Schulman, a co-founder of OpenAI who is now chief scientist at Thinking Machines, expects AI to make AI researchers ten times more productive in about two years. Charlie O’Neill, head of model training at Base 10, puts the same milestone five to ten years away. They gave these forecasts at the end of a Dwarkesh Podcast panel published September 11. Beren Millidge, chief technology officer of the open-source model developer Zyphra, was also on the panel.

The question defined "tenfold" concretely: if a breakthrough takes a year today, researchers would make one every month. The stakes come from a feedback loop. AI that speeds up the people building AI would speed up AI progress itself, a loop usually called recursive self-improvement. The rest of the conversation explains why two experienced researchers read that loop so differently.

Writing more code is not the same as discovering more

Schulman described a cycle he keeps seeing. A new model comes out, people are amazed and call it AGI, and then after about a month of use it "starts to feel dumb." Each new model catches up with humans in some areas. Research, he said, is held back by the places where the model is still weaker: where its judgment is worse, or where it cannot check its own work well enough. That is why, "even if the model can write way more code than a person, it doesn't make you like a hundred times more productive." He said it is hard to predict how many more times the cycle will repeat.

O’Neill's doubt was about paradigms, meaning the basic recipes the field builds on. He described a popular picture of takeoff: once an AI agent is even "0.1% better than all humans" at AI research, running hundreds of thousands or millions of fast copies would overwhelm every other bottleneck. He compared AI progress to Moore's law, the decades-long doubling of the number of transistors on a chip. That trend looked like a straight line, but many separate inventions were needed to keep it going. In the same way, he said, language models improved through pretraining (training a model to predict the next word across huge amounts of text) until the gains slowed. Reinforcement learning then kept the line climbing. It trains a model by rewarding successful attempts in practice tasks called environments. If the next jump needs a similar discontinuity, O’Neill said he is not sure that training models in RL environments, even ones aimed at self-improvement, would find it. The breakthrough might go deeper than adding a step to the current recipe. It might even mean giving up gradient descent or neural networks.

He separated two kinds of research. In one, the goal is already clearly specified, such as lowering a model's training loss (its error at predicting the next word), and the work is to optimize it. The other is open-ended science, where nobody can write the objective down, and paradigm shifts come from that kind of work. O’Neill said the AIs cannot specify that objective either.

Schulman offered an example from his own career. In the early days of OpenAI, he said, he believed that simply training a language model to predict the next word would not produce intelligence. The important information seemed to make up such a small part of the training signal that noise would drown it out. Humans don't learn to reproduce every detail of what they see, so he thought researchers needed better objectives. "But then it turned out that it just worked anyway." He said the biggest advances often come from generalization "that we have no right to expect." His examples: simple next-word prediction producing deep understanding, and skills trained on tasks with checkable answers carrying over to tasks without them.

A century of thinking around every experiment

The discussion then turned to a thought experiment. Before every seven-figure experiment, spend as much computing power on AI researchers as on the experiment itself. Automated researchers could spend the equivalent of a century planning the best experiment, running small test versions and building theory, then another century analyzing the results.

Schulman found the idea plausible. He said a clever enough small-scale experiment could probably have predicted some past surprises. "I would expect that like we're nowhere near the ceiling of how well you can do research," he said. He pictured AI spending about as much computing power on analysis and theory as labs spend on the experiments themselves.

O’Neill agreed only when the goal is well defined. Thinking can only update beliefs from information already gathered: "you can't gain any new bits from like just thinking." But when the objective is clear and the data already exists, he expected a large speed-up. His example was the 2020 scaling-law study by Jared Kaplan and colleagues. It measured how language-model performance improves with model size, data and computing power, and advised growing models faster than their training data. O’Neill said an AI would have noticed that the study used intermediate checkpoints (snapshots taken partway through training) without accounting for annealing, the gradual lowering of the learning rate at the end of training. By his estimate, that observation could have saved the field a year or two.

A 2024 study by Porian and colleagues supports part of this and complicates the rest. It traced the gap between Kaplan's estimates and later ones to three main factors: how the final layer's computation was counted, the length of the warm-up period, and tuning optimizer settings at each scale. Learning-rate decay improved results, but it was not needed to reproduce the later estimates. That study also used models below one billion parameters. O’Neill's second example was μP, a method from Greg Yang and colleagues that keeps training settings such as the learning rate valid as a network gets wider. Researchers can tune a small stand-in model and copy the settings to a large one. In one demonstration, a 6.7-billion-parameter model was tuned with a 40-million-parameter stand-in, at about 7% of the large model's pretraining cost.

That is where O’Neill's tenfold figure comes from. "I would imagine like a 10 times speed up if our thing is just like maximize the objective we're currently on," he said. But "just thinking doesn't necessarily buy you the right objective in the first place." He pointed to how the current paradigm began. DeepMind's plan was to solve intelligence by teaching systems to play games at superhuman levels. By contrast, Alec Radford, whom O’Neill called "one random researcher," tried predicting the next word across a very wide range of text. Radford later led OpenAI's 2019 GPT-2 report, which tested how many tasks that broad training could handle without task-specific training. Even then, O’Neill said, it took a while to scale the approach up, because the field first had to learn that scaling laws could reliably predict the results.

The job Schulman expects humans to keep

The host asked which part of AI research would be the last left to humans. The answer in the discussion was iteratively asking the right questions: today's AIs are much worse at deciding which experiments to run than at coding them, and when asked for research ideas they tend to suggest many very small steps.

Schulman said the human role that will last longest is "defining the objective, and like deciding what we actually want." Examples include deciding how an assistant should behave, what counts as helpful, and what goal to use when training from human feedback. He said the same goes for writing the constitutions and model specifications that labs now use to describe intended behavior. "Even if the AIs can do all the technical work, we'll have to still do a lot of that," he said.

He split alignment, the work of making AI do what people intend, into two jobs: deciding what the objective should be, and actually achieving it. The first, he said, is "not going to go away anytime soon." This is why post-training teams, which shape a trained model's behavior before release, need so many people. The model's behavior has to be worked out area by area, and someone has to decide how it should act in each one. That makes the whole job very hard to automate.

How an automated researcher might be trained

Asked how the first AIs able to automate AI research would be trained, Schulman described a practical mix. One part is learning from human feedback, so models absorb researchers' taste. The other is building many practice environments made of multi-step research projects. Each round would "patch" whatever seemed most broken last time: researchers use the AIs heavily, notice consistent weaknesses, and fix them with new feedback or new environments.

O’Neill described how this could work at a lab. Take the bugs a company such as Anthropic found in its training setup over the past few months and turn them into training environments. In principle, he said, a lab could go further back in its own history. For example, it could return to before GRPO, a reinforcement-learning method introduced in DeepSeekMath in 2024, and challenge models to rediscover the best way to train. But he expects labs to stay short of computing power and keep working at the frontier, turning each round's bugs and improvements into environments. He said this is essentially distillation, where a model learns to copy existing work, and it may be why progress can feel like it is levelling off. The models keep inching toward what human researchers have already found.

The conversation offered a counterpoint. Copying human work can never surpass it, but an environment can set a goal no human can reach, such as a lower training loss than anyone has achieved, and the model can keep trying. The idea resembles the community nanoGPT speedrun, which scores how quickly entrants can train a small language model to a fixed loss on eight H100 graphics chips, except that the AI would aim to beat the human speedrunners. Building a 100-million-parameter model that beats Minecraft was floated as another goal, then as perhaps too easy. "Isn't it crazy," O’Neill said, that a 100-million-parameter Minecraft player now counts as too easy, when five years ago it would have sounded remarkable.

Schulman said much of research looks nothing like climbing toward a clearly measured goal. It usually starts with an intuition that models should improve in some way, plus an algorithm idea that seems to point that way. Researchers then design a task meant to show "signs of life," relaxing realism for a while. If the method works, they make the task more realistic step by step. Other research aims to explain what is happening and build theory, usually informal because machine learning rarely has predictive mathematics. The discussion settled on the expectation that future models will train on all of these kinds of task: some checked automatically, some graded by another model or a human. Whether that training carries over well enough to fuzzy research judgment that humans could leave the loop entirely remained unclear.

What the timelines actually measure

The forecasts came in a rapid-fire round, and definitions shaped them. The first question asked when AI could work as a drop-in remote worker for all kinds of office work, handling complex, month-long projects with full computer use. O’Neill's answer ranged from about a year to a couple of years, depending on whether the AI must work through a browser or firms make their information easy for software to access. He added one thing he doesn't expect the model to do: really push a colleague to get something done. "It's just going to be too nice." Schulman said human remote workers vary widely in quality. On a freelance platform like Upwork, some work already falls below what current AI can deliver. He said he basically agreed with the others that some version would exist within a year, doing some things well and others poorly.

O’Neill said people keep moving the goalposts toward rare, difficult cases. This year, he said, he simply told OpenAI's Codex agent to collect everything he needed for his taxes and send it to his accountant. It clicked through a long list of tasks and downloads, and "it was fine it was perfect."

The tenfold research question was different. Schulman at first refused to give a single number, saying some kinds of work, such as certain mathematics, might already be past that point. Asked specifically about AI researchers pushing the field forward, O’Neill said five to ten years, and Schulman said about two. O’Neill's answer drew surprise, since it was much longer than his remote-worker forecast, and the conversation turned to whether the two questions had been defined differently. O’Neill said he was picturing normal white-collar work over about a month; beyond a month, he said, the two start to diverge. His crux was his own ability to absorb information and make the statistically best decision about the next experiment.

The case for the shorter timeline came up in the same exchange. The argument went that AI already speeds up coding by more than tenfold, so if it can run even two or three experiments in a row, each informed by the last, without failing, that would be a big uplift. Research would then be held back by other bottlenecks instead of by researchers' ability to run small experiments.

The final question asked about an AI better than top human experts at all computer-based work, including projects that take years. Schulman said three or four years. He said AI research gets a lot of attention and relies heavily on code and math, which models handle well. Fields involving 3D, spatial or physical work, such as mechanical engineering, might take longer. O’Neill again said five to ten years, agreeing that fully automating AI research is close to requiring superintelligence. Even with memory systems outside the model, notes to itself and somewhat longer context, he said, some things in the world still demand more context than a model can take in.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06

Connected ideas and articles

From the conversation

Podcast episodes

Dwarkesh Podcast

AI researchers debate how close we are to recursive self-improvement

Episode published This article draws on 0:00–17:23, 28:05–33:50 and 1:28:31–1:36:50 (approximate times)

Article history

Updates to this article

Tags