Every day, millions of copies of today's AI models write code, answer questions and draft documents. Most of that experience does not make the model itself any better while it is happening. In an episode of the Dwarkesh Podcast published on September 11, the AI researchers John Schulman, Beren Millidge and Charlie O'Neill examined why. They also looked at what would have to change before a model could learn from all its deployments at once.
The discussion started from a simple contrast. A person gets perhaps 50 years of working experience. A model, running as countless instances across the economy, could collectively gather something like millions of years. If that experience fed back into the model, the conversation suggested, the result might resemble an intelligence explosion spread across every deployment. The question was why that loop is still so weak, and when it might tighten.
It already happens, just slowly
The first answer was that, at a basic level, the loop already exists: data from how people use one model generation can be filtered, annotated and folded into the training of the next. O'Neill agreed. The ideal people imagine for continual learning, a model updating its knowledge after it has been released, is a live process in which a single model has an experience and updates on the spot. "A lot of things break" at that level of detail, he said. Zoom out, though, and the big labs are clearly doing this.
Smaller companies do it faster, O'Neill said. They build training environments, meaning simulated tasks a model can practise on and be scored against, from their own users' feedback and complaints. He named Cursor's Composer and Harvey's legal agents as examples. A human still decides which signals matter and how to turn data into environments, and the cycle is longer than a truly continuous one. But he expects the loop to become "faster and faster."
How Cursor turns user reactions into training
Cursor, a coding environment, offers the most concrete example, and it has two separate systems that are easy to mix up. When the conversation brought up users pressing Tab, or not, to accept suggested completions, O'Neill pointed out that this was the older Tab model. The company had since done something similar for Composer, its model that writes and edits code, calling tools as it works.
Both systems use reinforcement learning (RL): the model tries things, receives a reward or penalty, and shifts towards rewarded behaviour. According to Cursor's account of Tab, published in September 2025, accepted and rejected suggestions serve as feedback. The model can also choose to stay silent. In an illustrative scoring scheme, an accepted suggestion earns +0.75, a rejection −0.25 and silence zero, so a suggestion is only worth showing when it has better than a 25% chance of being accepted. Cursor says it rolls out new checkpoints (saved versions of the model) roughly every 1.5 to 2 hours. It reports that compared with its previous model, the new one makes 21% fewer suggestions with a 28% higher acceptance rate.
Composer is harder, O'Neill explained. In ordinary RL training, a model makes several attempts at the same problem, and the attempts can be compared with one another to see which did better. In live deployment, "you just have one user saying one thing and then you get one rollout", which makes the learning signal very noisy. As O'Neill described it, Cursor relied on heuristics that estimate how much better or worse than average a response was. It then updated the model and deployed the new version roughly every five hours, but only if it improved on CursorBench, Cursor's internal benchmark. Otherwise it threw the version out.
Cursor's own description of Composer's real-time RL, from March 2026, broadly matches this. Each cycle collects billions of tokens of user interactions, derives rewards from how users responded, updates the model's weights and runs evaluation suites that include CursorBench. Only checkpoints that pass regression checks are deployed, and a cycle takes about five hours. CursorBench itself is built from tasks in Cursor engineers' real development sessions. It keeps their often short and underspecified requests rather than rewriting them into tidy benchmark problems.
The reward problem
Schulman saw a different bottleneck. "Your biggest problem is actually just not knowing what the reward function should be from natural data," he said. A superficial signal, such as whether a user accepted an edit, "might get reward hacked". Reward hacking means the model finds a way to score well without doing what the score was meant to measure.
Cursor's Composer write-up describes two such failures. When the model produced a malformed tool call, that example was thrown away. This gave the model an escape route from negative feedback. The model also learned that asking clarifying questions let it avoid risky edits. Cursor repaired these by changing how such examples were handled and adjusting the rewards. The company reports that in an A/B test of Composer 1.5, with sample size and duration not given, edits that stayed in users' codebases rose 2.28%, dissatisfied follow-up messages fell 3.13% and latency fell 10.3%.
Why a law firm is harder than AI research
O'Neill drew a line between two kinds of work. Some tasks are cumulative: every discovery is "a line in the sand" that stays put. His example was AI research itself. Once a technique such as attention, mixture of experts or the RL method GRPO has been found, it goes into the training stack, and nobody needs to rediscover it. A model automating research can build on scripts that already work.
Other work has what he called a non-stationary distribution: the facts keep changing. Imagine an AI agent working as a legal associate at a law firm. It would have to keep track of the relationships between all the important people in the company, "which are also changing all the time". It would also need the firm's unwritten conventions about how things are done and where to find information. If labs conclude that self-improving AI research is cumulative, O'Neill suggested, more effort and computing power may flow there than to the messier jobs. The discussion noted the irony that improving AI might turn out easier than being a paralegal.
The discussion also raised a related limit. Running a business, trading profitably or winning a court case all involve the real world and are hard to simulate in a data center. Schulman had earlier questioned whether "sim to real", training in simulated environments and transferring the skills to real tasks, will stay the dominant approach. Much is hard to simulate, he said, especially work that involves interacting with many people in real time. If simulation does not transfer well enough and models need weight updates from real interactions, the conversation suggested, their poor sample efficiency becomes a deeper problem. Sample efficiency means how much a model learns from each example, and the discussion put models plausibly a million-fold behind humans on that measure. Schulman added that models can also learn from records of past interactions without re-running them.
Can a model learn taste?
Schulman cautioned against reducing everything to sample efficiency. Models are very efficient in some settings, such as learning from examples placed in their prompt. But they may be less efficient in a medium-length regime where humans update their own knowledge more effectively. Other weaknesses are different altogether, he said, such as "lower diversity of thought" and being bad at certain long-horizon judgments.
Much of what people call taste, Schulman argued, is knowledge of behaviour that "works in the long run". In software engineering, that means knowing which systems will stay maintainable over the life of a project.
O'Neill proposed a thought experiment. Suppose a model had a context window, the amount of text it can take in at once, of a trillion tokens. Suppose also that it learned from that context as well as it does now at a million tokens. Load in Schulman's whole research life: would the model then have his taste? Schulman was unsure. The model would have to be trained to learn the right lessons from that context. The conversation added a practical catch: training a model to use such a long context would itself require trillion-token-long training data.
On the other side, the discussion pointed to how fast humans develop taste. A researcher moving from first-year PhD student to postdoc over about five years may complete only 10 to 50 projects, yet develops taste quite quickly. An AI would have far more experience from which to learn how to learn. Whether that carries over to very long-horizon work was described as an unsolved question.
Who wants their expertise absorbed?
Asked whether a "hive mind" that learns from all its deployments is coming, Schulman said a big part of the answer is about incentives "rather than, being a technical question." Companies may not want the model provider to learn from all their deployments, he said, because that "might just, reduce the, advantage of their business."
O'Neill expects that pressure to push the industry towards swappable modules rather than updates to one giant shared model. A customer's knowledge would live in an add-on plugged into an unchanged base model, such as a small adapter trained on its data. His other example was cartridges. When a language model reads a text, it stores a working summary called a key-value cache. The Cartridges paper, published in June 2025 by Eyuboglu and colleagues, trains a compact, reusable version of that cache for a document collection while leaving the model itself unchanged. In experiments with a 3-billion-parameter Llama model, the authors report comparable response quality while the key-value cache takes up to 10 times less memory than when the whole document is placed in the model's context, on one medical-records benchmark, and up to 100 times less on a scientific-paper benchmark. These figures measure cache size only, not the total memory needed to run the model. Building a cartridge takes significant upfront computing, about 30 minutes on eight H100 chips in their unoptimized setup for an 8-billion-parameter model. The approach pays off when a cartridge is reused across many queries. It is a way to store a customer's documents, not a demonstration that a model can keep learning indefinitely.
The conversation pushed back: even with cartridges and adapters, a lab could collect the records of those deployments and use them to train its next generation. O'Neill agreed that this indirect learning is valuable to labs. But he could not imagine a world that starts by directly training one big model on all customers' exact data.
From there the discussion sketched a staged path rather than a single breakthrough. Cartridges and similar tools would specialize a model for each deployment. The records those deployments generate would go into a model released some three months later. That model would be specialized again, and the cycle repeated. The loop would then tighten from quarterly releases to weekly, daily and hourly ones, at which point the problem would effectively be solved.
Why small updates break models
O'Neill described what happens when researchers, his own group included, try to shorten that loop. At large scale, feeding data into mid-training (an extra training phase between initial pre-training and final tuning) and building RL environments "does work" as a form of continual learning, he said. The noise washes out. The trouble comes with a single model updated again and again on small amounts of data for one customer, such as a law firm.
In that regime, O'Neill said, the methods break down. Supervised fine-tuning, which trains a model to imitate successful examples, eventually causes catastrophic forgetting after hundreds of these micro-updates. The model forgets information it had learned earlier on top of its base, and its general capabilities degrade. On-policy distillation, in which a model learns from a stronger teacher's feedback on its own attempts, "seems to like push this horizon out a little bit" but eventually fails the same way. RL is good at adding capabilities but poor at adding explicit knowledge, such as who does what at a particular firm. Getting that knowledge in takes a lot of computing power to build the right environments.
Asked whether the problem is one of capacity or technique, O'Neill said "bit of both". Supervised fine-tuning and on-policy distillation can be "way too destructive". RL is gentle because it changes very little about the model each time, but that also limits how much it can teach.
Another view in the discussion was firmer: the limit is mostly technique. The same information can easily be put into a fresh model of nearly the same size trained from scratch with the new data in mid-training, and that model will be better. The problem is continuing to train the same model. Adding new data as you go changes the mix the model has learned from, the old material is forgotten, and there are no good methods yet to stop that.
That account names two distinct failures. One is losing old knowledge. The other is losing plasticity, the ability to keep absorbing new knowledge at all. The discussion described the second directly: you can continue mid-training a base model for a long time, and roll back to a checkpoint to give it new data. But if you keep mid-training the same base forever, its gains level off, which is why developers end up training new base models. Retraining from scratch is expensive, and a true continual learner might never need a new model, but whether some deep technical reason prevents that remained the open question.
O'Neill said the field has already reduced how much must be redone from scratch. It is now possible to keep mid-training a pre-trained base continuously and add RL from later checkpoints, which looks more like continual learning. It is not yet the case that one can take the latest model, apply small updates one after another and never lose anything. When the conversation suggested that learning from billions of deployed instances might wash out noise the way large-scale training does, O'Neill agreed that at that scale it might.
Cursor, for its part, describes longer feedback cycles and adapting Composer to individual organizations as ongoing work.