One of Applied Compute's customers runs an AI assistant that many people use. They ask it to do things, then tell it what it got right or wrong. Yash Patil, Applied Compute's CEO, said trying to capture all that feedback in written instructions, such as "some sort of giant skills file," is "pretty impractical." Instead, his company uses the feedback to change the model itself.
Patil, a former OpenAI researcher who worked on Codex, described the case on the Unsupervised Learning podcast in an episode published October 6, 2026. He said the approach had brought "real upticks in user satisfaction, completion rates, decreased tool call failures." He gave no figures. A tool call failure is when the model tries to use a piece of software, such as a search function or a database, and the attempt goes wrong.
The example points to a question many companies with an AI product now face. Should they keep refining the instructions and software around a model they already use, or train a model of their own? Patil gave his answer in steps during his conversation with Jacob Effron, an AI investor at Redpoint Ventures. Applied Compute sells exactly this kind of training, so his answer is also a pitch for his own company.
The two levers
A language model's behavior comes from its weights, the vast set of numbers it learns during training. Post-training means training a model further after its main training run, for example with reinforcement learning (RL), in which a model improves by trying tasks and being rewarded for good results. Companies can do this with open-weight models, whose weights are published for anyone to download and adapt.
The other lever leaves the weights alone. A team can write better prompts, choose what information goes into the model's context (the material it reads before answering), and build a better harness, the software loop and tools a model works inside. Patil called this optimizing by "configuration." Changing the weights optimizes the "policy," which he described as "how it does judgment, how it does reasoning."
His order of operations favors the cheaper lever first. "You're still better off going and doing a bunch of context optimization and harness optimization," Patil said. But once a team has "squeezed a lot out of the harness," he argued, it can still gain a lot by training the model "to use those tools better."
He gave an example from chips. One customer in the industry had built its own software, which general models "don't really know how to use well," for tasks such as ETL (extract, transform, load) verification. By applying RL on top of that specialized harness, he said, the company pushed models to "a lot better performance."
Capability or price
Effron noted that many early post-training projects he heard about were about cost and speed. Running similar tasks at huge scale makes the inference bill climb, and inference is the cost of running a model to produce answers. Customers also prefer faster products.
Patil agreed that this is the usual reason. Frontier models from the big labs still lead open-weight models on base capabilities, he said. A company can close that gap only with differentiated data, proprietary material that is out of distribution, meaning unlike what the model saw in training. "The more out of distribution your data is, the more likely that post-training is actually going to give you that capabilities lift," he said. Without such data, the gain is mostly economic: train a cheaper model to match the frontier on one task, then serve more users for the same money. "Most people are actually post-training to basically optimize price performance," Patil said.
What the labs can't buy
Effron asked how much company data will stay out of distribution as the big labs, now focused heavily on coding, move into other fields. Patil said the answer depends on how training itself changes. Today's dominant method, which he called offline RL, relies on high-quality datasets bought from data vendors. Training on those makes a model "very spiky" in one domain, he said.
What stays out of distribution, he argued, is data produced inside a company. His extreme example was an organization that makes its own data by running experiments in the real world and getting reward signals from them. In an ordinary business, he said, the equivalent is "judgment traces": records of decisions shaped by a company's thresholds for risk, operating model, expertise and historical data. Two companies can face the same situation and decide differently.
Patil expects online RL, in which a model improves from its own production use, to become the way enterprises capture that judgment. "The more they use the model, the better that it gets," he said. He predicted that in five years many companies will have systems that gather input from experts and turn it into better policies.
Applied Compute's written material makes a similar case. In a June 2026 essay, Sahar Solimano described user corrections, retries, edits and overrides as learning signals, with domain experts defining what good work looks like. The company's AC2 platform announcement from August 25, 2026 describes feeding production interactions and feedback into later training runs. The product was announced as a private beta.
Which tasks need changing weights
Effron pushed back. Anyone who has played with these tools knows that strong prompt engineering or well-made skills can carry a lot of knowledge, he said, and researchers can be "a little bit of like a hammer looking for a nail." How confident was Patil that companies would need many models rather than other ways of stitching things together?
Patil said his company is "pretty confident that the future looks like a non-static set of weights," with some feedback loop from the real world back into the model. Then he asked whether that is needed for every task, and drew a line. Things that change over time call for a closed loop in which new information makes its way back into the model. The tasks that need iterative training are the judgment-oriented ones. Manual, process-oriented tasks, he said, may eventually be covered by the "expanding umbrella of intelligence" that base models provide out of the box.
The company's own essay leaves similar room. Solimano wrote that many systems should keep using general models with harnesses, context, tools and retrieval, and that customization matters most for workflows central to product quality, margins and differentiation.
Patil also said Applied Compute is not tied to any single training method. RL and online methods are "the best things right now," he said, which is why much of the team works on research systems, testing new methods in production with customers whose deployments serve "tens of thousands, hundreds of thousands of users."
Making training cheaper
For Patil, the deciding question is the total cost of ownership: "how expensive is it to post-train?" Much of his company's work, he said, aims to lower the cost of applying RL to a model, and to do it automatically and online, without customers building hand-tuned RL environments that are fully replayable. An RL environment is the task, tools and grader a model practices on. The AC2 announcement says the company's self-distillation workflow can use production traces even when the original environment cannot be replayed.
Patil said the cost of post-training is falling, and he pointed to three levers. First, hardware: "you'd be surprised how underutilized GPUs are," he said, and more can be squeezed out of existing chips. He credited his co-founder and chief architect, whom the company's founding announcement names as Linden and credits with RL training infrastructure work at OpenAI. Second, algorithms that learn more from less data, so training runs are shorter and use less compute. Third, synthetic data, which can take some amount of tokens and "explode it into something much larger." He expects most of the savings to come from systems and algorithms.
Other companies are selling similar plumbing. On August 31, 2026, Fireworks made its Training API generally available. It lets customers control their own training loop, data and rewards while Fireworks runs the computation and moves finished checkpoints into production serving.
A model for every firm?
Effron cited Microsoft CEO Satya Nadella's line that there should be as many models in the world as there are firms. Applied Compute's homepage features the remark. Effron called it "a good selling point for you all" and asked whether Patil believed every company is a special snowflake that needs its own model.
Patil said companies do not necessarily need to pre-train or mid-train their own models, the costly early stages that build a model's general knowledge. But each company has closed-loop feedback systems worth capturing, he said, even if most are starting with harness and context work because there is still "a lot of low-hanging fruit." In his view, great shared models "set the floor for everybody," while how companies optimize those models and build systems around them "is what's going to set the ceiling."
Effron offered a range of cases. A pharmaceutical company, whose experimental data in one disease area is unlike anyone else's, should obviously have its own model, he said. But how different is one bank from another, if a frontier lab trains on plenty of good general finance data?
Patil split the question. Fragmented businesses such as accounting, where one accountant serves one client, are ripe for consolidation, he said: take the routine labor out of an accountant's work and put it into a model. Banks are different, he argued, because each makes "a whole litany of decisions" and trade-offs that the others don't. That, he suggested, is why some banks do well and others do poorly. "There actually is, like, room to compete," Patil said.