October 6, 2026
Heard in AI

Yash Patil says cost is driving AI model choice more than he expected

Applied Compute CEO Yash Patil, whose company trains custom models, says firms now weigh cost against performance instead of chasing the biggest model. He cites OpenAI's Baseten deal, OpenRouter usage and cheap task models like TypeSafe's Jev.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Unsupervised Learning, episode published October 6, 2026

A couple of months ago, Yash Patil said, the AI industry wanted to "promote and analyze the strongest, biggest, baddest model out there." Now the conversation is "much more" about a curve: how much performance a company gets for what it pays.

Patil is a former OpenAI researcher who worked on Codex, the company's coding agent. He now runs Applied Compute, which specializes models for particular companies through post-training, meaning additional training on top of an existing model. He made the case on an episode of the Unsupervised Learning podcast published October 6, 2026, hosted by Redpoint Ventures investor Jacob Effron. His company profits if customers decide that cheaper, specialized models are worth having.

The deal that prompted it

Effron opened the subject with a recent partnership between OpenAI and Baseten, a company that runs AI models for customers as a hosted service. In Effron's description, enterprises could use their OpenAI credits to "burn them down" on open models served by Baseten. Baseten's September 29 announcement bears that out. Baseten became an early provider of open models in OpenAI's business marketplace, and enterprise customers can apply their existing OpenAI commitments to those models through Codex or OpenAI's Responses API. Open models are ones whose weights, the learned parameters that make up a model, are published, so companies other than the original developer can run them.

Patil called himself an admirer of the Baseten team and read the deal as "a clear signal" that "it's a multi-model future. Customers want choice. They want flexibility." For OpenAI, he said, it shows a move toward being a platform: "We want to be the gateway to enterprises using AI." He also called it "a big win for open source models."

Effron asked whether OpenAI could really act as a neutral broker that offers enterprises whichever model is best. Patil said Applied Compute was excited about taking part, but saw a difficulty. A marketplace makes obvious sense for Amazon, he said, because its Bedrock service already runs Anthropic models, OpenAI models and open models side by side. Whether OpenAI would go as far as offering Anthropic or Google's Gemini models, he said, "I'm not sure." Amazon's announcements describe the Bedrock lineup: OpenAI models joined providers including Anthropic, Meta, Mistral and Cohere in a limited preview in April 2026, and OpenAI's GPT-6 Astra became generally available there in September.

From the biggest model to a Pareto curve

The curve Patil described is a Pareto curve. Each point on it is the best performance available at a given price, so getting more capability along it means paying more. "You don't need a mega model for every single task," he said.

Open models, in his account, give buyers more points to choose from: "you can slide along the Pareto curve by picking different open models of different sizes from different families." Applied Compute's "core thesis," he said, goes one step further. Training a model to be better at particular tasks can move the whole curve outward, producing "better models for cheaper cost."

He linked the shift to where companies are in adopting AI. Businesses have grown comfortable running AI in production and have done a lot of exploring, he said. The mood now is: "this AI thing is really working. Let's go ahead and optimize it."

Later in the conversation, Patil tied this to scarce computing power. "We are in a supply crunch," he said; compute is limited and "everyone's kind of feeling this." The more a company can squeeze out of its models and the hardware it has, the more it can do. That holds, he said, even when a cheaper model pushes the curve outward without advancing the frontier of what AI can do. "If I can do something 10x cheaper than my competitor, that is differentiation."

The paradox he underrated

Asked what he had changed his mind about in the past year, Patil said he had not fully appreciated "how real" Jevons paradox is. The idea goes back to the 19th-century economist William Stanley Jevons, who observed that making steam engines burn less coal per unit of work made many more uses worth pursuing, so total coal demand rose. Applied to AI, cheaper answers can mean more total spending on AI, not less.

Patil's evidence was OpenRouter, a service that sends developers' requests to many different models. In its statistics, he said, "whenever there's a price cut for a model, usage just spikes." A usage study by OpenRouter and the venture firm a16z, published in December 2025 and covering more than 100 trillion tokens, gives a more qualified picture. (A token is a small chunk of text, the unit in which AI use is metered and billed.) It found heavy use of inexpensive models, but also substantial demand for expensive ones, and only a weak overall link between price and usage. The authors read some adoption of cheaper models as Jevons-like: people using longer inputs, more attempts and new applications. But the analysis observes usage rather than testing cause and effect, stresses quality, reliability and fit for the job alongside price, and covers only traffic that passes through OpenRouter.

Patil still drew a broad lesson. "We're moving to a world where things are, like, extremely cost-optimized," he said. Everyone is under pressure "to, like, make things as efficient as possible, lower the AI bill." He had not expected cost to matter so much compared with capabilities. Most open models, he said, can now "kind of do everything that you might want them to do," though he stopped himself: there are many frontier tasks, "I don't want that to be a blanket statement," and "Frontier models are amazing." What surprised him was how much people care about "using the right model for the right task."

Effron noted that companies had used frontier models for some tasks because they got real value from them. Patil described a sequence: there is some marginal return on extra capability, and at some point people will experiment, "try a bunch of things, and then consolidate on the right model."

A model named for the paradox

That brought the conversation to Jev, a model from the startup TypeSafe. The company announced it in early access on September 15, 2026, and its founder, Diogo Almeida, said it is named after Jevons. Jev does not write free-form text. Software gives it a situation and a fixed list of possible answers, and it returns decisions with probabilities attached, producing its outputs in parallel. TypeSafe's announced price was $0.042 per million input tokens, with no charge for output. Suggested uses include classification, routing requests and evaluating records of what other models did.

"It's in the name," Patil said: things that are "cheap, fast, designed for, like, a particular set of tasks." He said that is how Applied Compute thinks about much of what it builds, and he expects more models like it. His company is working out how to use Jev inside training for "massive billion token scale classification on rollouts." Rollouts are the attempts a model makes during reinforcement learning, a kind of training in which a model tries tasks and is rewarded for good results. Work like that had been cost prohibitive, he said, and it was out of reach even with "the smallest, cheapest" model from the Qwen family.

Patil described Jev as "much less of a replacement for stuff rather than an enabler." He meant jobs Applied Compute could not do before, in training and in running models, such as observability: monitoring what deployed models are doing. Because these gains stack, he called the effect "highly multiplicative."

When Effron asked for a concrete example, Patil pointed to a blog post in which his company used Jev for "massive trace analysis for online training." A trace is the step-by-step record of an AI agent's attempt at a task. In the study, dated September 23, Applied Compute's Bryan Lee took 148 published traces from OpenAI's GPT-5.2 working through 69 banking scenarios in a benchmark called τ³-bench. He sorted the failures into 14 categories and compared how well, and how cheaply, different models labeled them.

The results depended on the setting. At a decision threshold of 0.5, Jev was the cheapest option, but a model called Luna, at medium settings, had the highest micro F1, a combined accuracy score. When the threshold was lowered to 0.20, Jev reached 85% recall and gave the best micro F1 for its cost. Extrapolating from this dataset to one labeling pass over 10,000 traces, the study estimated costs of $11 for Jev, $54 for Luna and $479 for Anthropic's Haiku 4.5, which is more than 40 times Jev's cost. Jev's calibration error, a measure of how well its probabilities match how often it is right, was 0.051 against Luna's 0.154. There was a practical limit as well. The Jev version tested accepts at most 32,000 tokens, so long traces had to be summarized or split before it could read them. This is Applied Compute's own study of one banking dataset, not an independent benchmark.

The point of the labels, in Applied Compute's setup, is what happens next. Each annotation is attached to the original trace, so engineers can open the failures it flags and decide what the model should be trained on next.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Connected ideas and articles

Background

Stories behind this article

  • TypeSafe's Jev launch and the debate over whether it is novel The article is mainly about cost versus performance in model choice. Jev is the example Patil gives of cheap, task-specific models, and it gets a long section on what it does, its price and Applied Compute's trace-analysis study. It does not cover the launch debate about whether Jev is novel or trivial.

From the conversation

Podcast episodes

Unsupervised Learning

Ep 94: Applied Compute CEO on the Limits of RL, the New AI Hyperscaler & Why Post-Training Wins Inference

Episode published This article draws on 4:35–7:19, 23:12–24:06 and 42:41–46:17 (approximate times)

Article history

Updates to this article

Tags