The discussion started with a question: why haven't AI model providers consolidated into a handful of dominant companies? Building a frontier model takes large amounts of computing power, data and money, and all of that seems to favor a few winners. On a panel for the Dwarkesh Podcast, published on September 11, 2026, AI researcher John Schulman named the main force pushing the other way: distillation.
Distillation means training one model, the "student", to imitate the outputs of another, the "teacher". Schulman's argument turns on reinforcement learning (RL), in which a model tries tasks and is rewarded when it succeeds. "Anything that can be learned through RL can be distilled very easily," he said, because the behavior RL adds amounts to "a small number of bits". If you can collect trajectories that show a behavior, meaning records of the model working through tasks, you can copy that behavior from a fairly small amount of data.
The catch: which questions to ask
The conversation quickly reached the complication. To copy a model's behavior, you first have to give it the right requests. Schulman agreed and went further. When distilling with supervised learning, he said, the prompt distribution is "extremely important". That is the range and mix of requests used to draw out the teacher's answers. Distilling all of a model's useful capabilities is "very non-trivial", he said, "even if you have full access to it" and can see its chain of thought, the step-by-step reasoning it writes out before answering. You need "a really wide distribution of realistic prompts".
That is why Schulman pointed to a possible shortcut that some Chinese companies may be using. He said US frontier models would otherwise be blocked in China, but router or proxy services let people there use them, mostly for coding. According to Schulman, some of these services collect and sell the resulting data. That would give a distiller "the perfect prompt distribution": real requests from real users. He described this as something that had been "coming out recently" and said the companies were "probably" using the services. He did not present it as a confirmed finding.
Can models write their own prompts?
The panel disagreed about how much this bottleneck matters. One argument was that AI already reduces it. Published pipelines, including those described in Chinese labs' papers, start with seed prompts from humans and other sources. They then use existing models to expand those seeds into broad coverage, so humans need to supply less information as the models improve.
The counterargument was that real use is messier than a list of tasks. A user asks for an application, finds it doesn't work, asks for a new feature, then decides to step back and try something else. Capturing that whole trace is the hard part. The reply was that if a model could invent such traces without help, you would already have recursive self-improvement: AI choosing its own training data and running its own training.
The discussion then tested a harder case. What if you wanted a model to be a really good politician, anticipating from scratch how a debate in the Senate might go? The answer was that this is, ironically, easier for distillers than for frontier labs. A distiller can simply ask a frontier model that already knows how to play the politician and have it produce endless variations. The first lab to build that skill has to find data on what politicians actually do each day. Reproducing a capability someone else has already built is much cheaper than creating it.
A puzzle about Anthropic's smaller models
Charlie O'Neill, another researcher on the panel, raised a puzzle that he thought undercut the prompt-data explanation. In his assessment, Anthropic's Sonnet 5 and Opus 5 are "almost objectively worse models" than the Chinese models GLM 5.3 and Kimi K3. He noted that Anthropic's models had access not only to ordinary distillation but also to logit distillation from Mythos, a larger model. In logit distillation, a student learns from the teacher's full probability estimates for each next word, not just its final text. O'Neill said the best measure of a capability is the very hard RL training environments built at the frontier. If Anthropic had those environments and that teacher and still produced a worse model, he suggested, frontier labs may have little advantage in RL environments now, if any.
He offered a possible explanation: an "uncanny valley" where the gap between student and teacher is too large. He also passed on a point others had made about Opus 5 compared with Opus 4.6. Opus 5 seems to have an AI judge checking everything it has done, which is why it uses so many tokens. Yet it lacks the "big model smell" of Fable, the larger model, which tells it when to stop checking and which path is worth pursuing.
Schulman's answer: difficulty and realism are different
Schulman proposed what he called "a slightly different hypothesis". Training environments, he said, vary along two axes. One is difficulty. It is comparatively easy to create many hard environments: complicated tasks or puzzles that demand cleverness and are easy to check. He called this the benchmarking distribution, because many prominent benchmarks look like that. The other is realism, such as a coding agent that works through several rounds of back-and-forth with a human and has to balance several objectives at once.
A lab building a behavior for the first time has to push on both, Schulman said. Getting good behavior on the realism axis takes rubrics or some kind of human feedback to shape the reward. Naive distillation tends to match the teacher only on the "bench maxing" distribution. Without enough environments that exercise the trickier, realistic cases, the student never picks those capabilities up. He added that one possible factor is that big models generalize better from narrow, hard tasks to realistic ones. So a student trained only on easily verifiable tasks can match the teacher on the benchmarks and still do worse across the broader range of real use. With a good, realistic prompt set, he said, a student can match the big model very well.
Schulman suggested this "might even explain something" about smaller Anthropic models like Sonnet 5. He added that it is hard to know what Anthropic does in post-training, the stage of training that shapes a model's behavior after its initial training on large amounts of text. The company might be constantly changing its post-training setup and have got a few things wrong, turning something up too high and creating quirks people dislike. "It's really easy to screw up post-training," he said, in ways that don't show up in benchmarks.
Other ways to keep up
Two more points in the exchange suggested why the leaders' head start may be thin. First, the discussion noted that frontier labs buy their data from data companies, and Chinese labs can buy the same data. "And they are," O'Neill said. Some people are annoyed by this, the discussion continued, but the same purchased data combined with distillation makes it fairly easy to keep up.
Second, Schulman said the picture could shift if companies learn to keep improving their own models from deployment, which he said would "change the game a bit". The response was that continual learning does not stop distillation. A model that improves every day could also be distilled every day, so the copying could keep pace with the improvement.