October 6, 2026
Heard in AI

Applied Compute's CEO argues post-training wins inference

Yash Patil, CEO of Applied Compute, argues that the biggest AI workloads are the most valuable both to custom-train and to serve, which he says gives a training provider an edge in inference. He calls serving the easier job.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Unsupervised Learning, episode published October 6, 2026

Yash Patil, chief executive of the startup Applied Compute, says most of the money spent on models goes into running them, not building them. He thinks that is the reason a company that trains custom models should also be the one serving them. "Post-training actually wins inference," he said on the Unsupervised Learning podcast, in an episode published October 6, 2026.

Patil is a former OpenAI researcher who worked on Codex, OpenAI's coding agent, before co-founding Applied Compute, according to the company's founding announcement. He was talking with Jacob Effron, an AI investor at Redpoint Ventures who hosts the show. Patil is selling this argument to customers, so it is his company's pitch as much as an industry observation.

Training, post-training and inference

The argument rests on three terms. Training is how a model learns, by adjusting its internal numbers, called weights. Post-training is a later round of that adjustment: a company takes a finished model, often one whose weights are openly available, and tunes it further for a particular job. Inference is running the model to produce answers, measured in tokens, the word fragments a model reads and writes.

Patil said Applied Compute treats post-training as "the differentiation and the part that makes the business sticky", meaning the work customers cannot easily take elsewhere. But he was clear about where the money goes. "If you look at the dollars spent on models, the vast majority goes to actually serving and using the models," he said. He said he thinks inference is going to be one of the biggest markets in the world, "if not the biggest", and added that this was not an unusual view.

Why the biggest workloads come first

Patil described a pattern he has seen in customers. Companies start with frontier models, the most capable systems from the big labs, and experiment. Once a use case matures and scales up, they look at open models to cut costs and get better performance for the price.

The workloads where companies spend the most are the ones they will want to post-train first, he said, "because that's actually where the most gains are to be had". He described spending across a company's use cases as following a power law: a few very large workloads account for most of the volume. If he is right, the same workloads are the most valuable to train and the most valuable to serve.

Patil then argued that owning both lets a provider tune them together. How a model is trained shapes how its serving should be set up. A long-running agent that calls many tools needs a different serving setup than a workload that mostly generates long outputs. He mentioned "disaggregated" setups that split the two phases of inference across different types of chips. The first phase, prefill, processes the input. The second, decode, writes the answer one token at a time. When Effron asked what this looked like in practice, Patil described it in general terms and did not give a specific configuration. The open-source serving framework SGLang lists separating prefill and decoding among its optimizations, so the technique is not unique to Applied Compute.

The history question

Effron pushed back with some history. Early open-source model developers said that building a model helped them serve it more efficiently. In practice, he said, a serving company such as Fireworks would thank them for training a great open model and then run it as efficiently as possible itself.

Patil answered in two parts. First, the low-level code that runs on the chips, known as kernels, should match between training and serving, and much of Applied Compute's work on training speed carries over to inference. Others make a similar case. Fireworks, launching its Training API on August 31, 2026, stressed matching numerical formats and kernels between training and sampling. Applied Compute's own AC2 announcement from August 25, 2026 says its serving keeps the sampling settings, numerical precision and kernels used in training. AC2 was announced as a private beta.

Second, Patil said there are two ways to lower an inference bill. A company could optimize its serving and cut costs by 10%. Or, because he thinks many models are "quite token inefficient", Applied Compute could train the model to use 10% fewer tokens while scoring the same on an evaluation the customer cares about. He tells customers the two are "two sides of the same coin", which he said is why the company wants to be the single provider for both training and inference.

Serving as the easier job

Effron asked how hard it is to build a good inference engine. Patil's answer undercut the idea that serving is a moat, a lasting advantage over competitors. The main requirement is capacity: the GPUs to run models on. He said open-source projects like vLLM and SGLang are good starting points. In his account, almost all inference clouds have retired many of their own custom engines in favor of one of these, then tune it for their workloads and add a few custom kernels.

He pointed to a remark by Dylan Patel on another podcast: put vLLM on some GPUs, list it on OpenRouter, a service that sends developers' requests to different model providers, and you roughly have a production inference offering. Getting listed takes more than that, though. OpenRouter's provider page describes a technical review and test traffic before listing. It also says applications are backlogged and that providers of proprietary models are being prioritized.

Patil did not want to "trivialize it too much". Faster token speeds and efficiency take a lot of work, and that matters at scale, he said, because "1% of savings across tons and tons of chips actually adds up." But getting production-quality serving running is "definitely the easier thing to do" of the two.

An inverted pyramid

Effron asked how an infrastructure company decides what to build and what to leave alone, since many seem to expand into each other's territory. Patil described his long-term aim: "a new AI hyperscaler." Hyperscalers are the giant cloud providers. In his telling, AWS, Google Cloud and Azure took CPU computing and built networking, storage, databases and application tools on top. The new one would be built on GPUs and on models, which he described as "this weird thing" made of weights and parameters. He called the inference engine the first piece of software to achieve real stickiness in the AI stack.

He described a series of layers. Training turns GPU compute into "more intelligent tokens". Inference serves those tokens. Routing picks which tokens to serve. Beyond those are security monitoring, identity and sandboxing, which keeps agents in a contained environment. He called post-training "a massively underlooked part of the stack".

He pictured this as an inverted pyramid. At the bottom are things very few companies can do: model training, which needs robust infrastructure and a GPU fleet you own and operate. Above that come inference, then routing, then context and harness, the software loop around a model that lets it carry out tasks. Many more companies can do the work near the top. Applied Compute started at the hard end, Patil said, and plans to widen its offering as it matures. "We're 16 months old right now," he said, and the first step was moving from training into inference.

For now, the company stays at the model layer. Effron suggested that many companies approaching Applied Compute could probably solve their problem with a better harness and some data work. Patil said its customers already have mature AI products and strong product teams building their own harnesses. Applied Compute does not focus on harness or context work, he said, and does not build agents from scratch. "We optimize the models behind mature products."

Picking customers

That focus shapes who Applied Compute works with. Patil said the company goes after very high-value use cases of three kinds. In some, a capability gain is worth a lot, as at a pharmaceutical, cybersecurity or chip company. In others, the inference workload is huge: billions or trillions of tokens, where a small gain in quality or efficiency adds up across a large user base. He said these are "typically the best customers." The third kind involves very unusual data that a customer wants to train on because it is unique to them.

Effron asked how a fast-growing company budgets for compute. Patil said the aim is "a very sticky, differentiated business", so Applied Compute is not chasing every customer who just needs inference. Instead of looking at total demand, he said, it picks the most strategic companies it can grow with: the ones that can prove his claim that custom models are stickier and more valuable than off-the-shelf ones, and that their returns last.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06

Connected ideas and articles

From the conversation

Podcast episodes

Unsupervised Learning

Ep 94: Applied Compute CEO on the Limits of RL, the New AI Hyperscaler & Why Post-Training Wins Inference

Episode published This article draws on 27:01–38:22 and 46:17–47:19 (approximate times)

Article history

Updates to this article

Tags