Yash Patil does not think the big AI labs are out to steal anyone's data. He still thinks companies should be wary of depending on them.
Patil is a former OpenAI researcher who worked on Codex, OpenAI's coding agent. He now runs Applied Compute, a startup that helps companies build and train their own AI models. In an episode of Unsupervised Learning published October 6, 2026, the host, Redpoint Ventures investor Jacob Effron, asked him which arguments for "owning your own intelligence" hold up. Effron listed the familiar worries about relying on frontier labs, the companies that build the most capable general models. Labs might withdraw access to their models. They might decide that a use is a safety risk. They might build products that compete with their customers, or train on customers' data.
Patil rejected the version of the case that "presumed that the labs are malicious." He drew on his time at OpenAI. The labs, he said, are "made up of great people." They have contractual obligations, and "a lot of engineering" goes into protecting customer data. As he put it, owning your intelligence "really is about flexibility and control."
What control means in practice
Patil said the idea gained traction when open models started getting good enough to do real work with them. For him, though, control goes deeper than choosing an open model. It means deciding where a model runs, how it is deployed and what it is optimized for, "whether it's cost or latency or specific domains."
That has concrete requirements. Teaching a model to do a company's tasks on the company's own data means changing its weights, the billions of numerical settings that hold what the model has learned. So, Patil said, a company needs access to those weights. It also needs infrastructure to "post-train" the model, meaning further training after a lab releases it. And it needs the systems to run the model for users and to route requests among several models.
That is Applied Compute's business, so Patil is describing what his company sells. Its platform overview describes customers deciding what gets trained and which versions are deployed. It also describes letting them switch the underlying base model without rebuilding the rest of their system. In its founding announcement, the company said customers own the specialized agents it helps build.
When the supply gets cut
Patil said there are real risks in being "super dependent on a single model vendor." The labs, he said, are doing their best to serve as many customers as possible. But models have been pulled from products "more than once, particularly in the coding space." As labs move into specific industries, he added, "it's not clear whether they'll continue serving companies that they're directly competing against."
Later, Effron asked whether such cutoffs would become more common. Patil named Anthropic and the coding tool Windsurf and said "it happened twice." "Competitive dynamics are real," he said. As labs build applications that compete with their largest customers, "I wouldn't be surprised if it happened more."
The Windsurf episode is documented. In a June 3, 2025 statement, Windsurf said Anthropic had given it less than a week's notice before removing nearly all of its direct capacity for several Claude models. Windsurf removed direct Claude access for free users, let people bring their own API keys and temporarily discounted Google's Gemini 2.5 Pro. A July 16 update said paid users had regained direct access to Claude Sonnet 4. Patil did not identify the second incident.
His conclusion was that the "most durable" choice is to invest in the full stack needed to build the models behind a company's products, without "extreme vendor dependence on any one set of models or companies."
What makes a data vendor worth paying
The same question of where expertise comes from runs through the companies that supply AI training data. Effron noted that in a lecture at Stanford, Patil had said he was "short" general data vendors such as Mercor and Surge. In investing, being short means betting that something will fall.
Patil said he was "not short the companies themselves." He called them great companies and said Applied Compute counts many of them as customers. His point, he said, was that more and more training data is synthetic, meaning generated by AI models. Many companies that build RL environments use models to create data that is then fed back into models, which he called "a little bit of a circular motion." RL environments are simulated tasks with tools and scoring rules, used in reinforcement learning, where a model learns from rewards.
The bigger issue for these vendors, he argued, is that research fashions keep changing. He traced the sequence. First came supervised fine-tuning on example answers and preference tuning. Then came RLHF, reinforcement learning from human feedback. Then came RLVR, which uses rewards that can be checked automatically. "Now the big thing is, like, robotics data." For data companies, he said, "the main thing to get really good at is pivoting."
Asked what sets one RL environment company apart from another, Patil said "it really is the human part": the expert networks run by companies such as Mercor and Handshake. At OpenAI, he said, his team worked closely with Mercor on early agent products, including Deep Research and Codex. Those networks reach "the best of the best in the world for coding or math."
The reason quality matters so much, he said, is scale. Training on a task uses relatively few examples, "a few hundred to a couple thousand per task or something like that." With so few examples, weak quality control has a "material effect" on the final model. The vendor that wins can quickly run a campaign across a wide network of experts and narrow the results to the best data points. The vendors present themselves in similar terms. Mercor says it vets domain experts to build tests and scoring rubrics. Handshake says it recruits academic experts through universities and verifies their credentials.
Why the bankers' edge stays inside the bank
Effron pushed the question further. Mercor can recruit former bankers to design training tasks. J.P. Morgan has its own bankers, with their own tasks and the tests they create to judge good work. He said he did not know how big the gap between the two would turn out to be.
Patil said bankers who leave lose access to "a lot of the proprietary data that's been built up over time." Their reasoning and judgment, he said, is "very heavily informed by the data that they look at." He called expertise seeping through a company's walls "a real source of, like, alpha," using the investor term for an edge over competitors. Companies, he noted, have long used non-compete agreements to guard against that.
Effron raised the practice of buying failed companies for their data. If he ran a bank and a rival collapsed, he said, he would not want one of the labs to acquire it. "Exactly," Patil replied. To him, that showed why such data is valuable and why "the more you can keep things sort of, like, proprietary to you, the better."
Effron wondered whether the labs would eventually work out the basics of running a bank. Patil compared running a bank to a cake with a thousand layers. He said it was unclear whether a single model, even a superintelligent one, could do that "out of the box." People have built up years of expertise in these jobs, he said, and "codifying those into models is actually, like, quite difficult." He did not suggest the jobs were safe, though. "There will be large portions of roles at different companies that will get automated away, for sure."