1 October 2026
Heard in AI

OpenAI's Ari Weinstein on what to hand computer-use agents

OpenAI's Ari Weinstein says computer-use agents can now operate software made for people, citing a two-hour meal order an agent did in 15 minutes. He adds that builders must earn trust through reliability, consent before payments and site limits.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published 30 September 2026

Ari Weinstein's meal-prep service lets him specify an order down to the gram: this many grams of chicken, this many grams of rice. That precision had a cost. By his account, a single order took him two hours. When he gave the job to OpenAI's computer-use agent running on the new GPT-6.1 Sol model, it finished in 15 minutes. He described that as eight times faster, which matches his figures.

Weinstein leads product and engineering for computer use at OpenAI. He co-founded Sky, a computer-automation startup that OpenAI acquired. He told the meal-prep story on Latent Space. The episode was recorded at OpenAI's DevDay and published on September 30, 2026. "Computer use" means an AI model operating software the way a person does: reading the screen, clicking and typing. The alternative is working only through APIs, the formal connections that apps offer to other programs. The conversation focused on what people are already handing to these agents and what builders owe the people who trust them.

A computer of its own

One of the day's announcements was Dots, which OpenAI introduced as always-on assistants. Weinstein said each Dot gets "its own Linux virtual computer in the cloud." He said this sets Dots apart from OpenAI's earlier products, whose agents could use either a browser in the cloud or the user's own computer. With a whole machine, a Dot can run full desktop applications as well as a web browser. According to OpenAI's announcement, users can inspect that computer and connect their applications. Background research that a Dot does without being asked is limited to read-only tools, which cannot send messages or control a computer.

Weinstein's case for the approach is simple. "All the software in the world was designed for humans," he said, and agents can now use that same software, so people can delegate to them. His advice for getting started was to pick something you already spend time on and ask whether an agent could do it.

Errands without an API

Flight booking, shopping and even playing games came up as candidates. So did YouTube. The discussion described features such as A/B testing and community posts as unavailable through YouTube's API, so the workaround was to run the agent in a virtual machine and let it use the site directly. YouTube's help page describes its title and thumbnail testing as a workflow in desktop YouTube Studio. It tests up to three variants and picks a winner by watch time. Weinstein said OpenAI's own developer experience team uses computer use with YouTube a lot.

Customer-service chats were another example. Replies there can take anywhere from 30 seconds to three minutes. The discussion described using Codex to deal with support bots, with one catch: the agent was too polished. It wrote complete sentences, capitalized correctly and gave full reference numbers, which made it obvious that a bot was typing. The fix was to prompt it not to pretend to be a bot but to act like a very annoyed human writing short one-liners. While it waited for replies, it was told to use subagents (helper agents it can spawn) to research better ways to get what was needed. Weinstein responded that he feels half the time it is a bot on the other end, "so now you've got the bots talking to each other."

The conversation then turned to how much trust has already grown. Three or four years ago, people were wary of connecting language models to the web and their devices. One account in the discussion described letting computer use configure DNS, the settings that point a domain name at a server, and pay bills worth tens of thousands of dollars, with a shrug about the worst that could happen. That led to the question for Weinstein: what should developers watch out for now that this capability is available through OpenAI's API?

Trust has to be earned

Weinstein first made the case for building on OpenAI's own setup, which is now available in the Agents API. Computer use has a universality to it, he said: it works with any website or service. Developers can build their own harness, the software loop that turns a model's decisions into clicks and keystrokes, but he said that is hard. He also noted that OpenAI trains its models on its own computer-use harness. Using the version the model already knows "might" bring advantages in speed, cost and accuracy, he said, offering that as a possibility rather than a measured result.

He then turned to trust. People are still getting comfortable with the technology, he said, and some may be further ahead than most. "It's incumbent on us to sort of build that trust over time," he said. In his view, that means building reliable products and the right safety checks. It means asking for the user's consent before doing something consequential, such as making a payment. Depending on the application, it also means letting the agent reach only the websites or applications the task actually needs.

OpenAI's Agents API documentation for computer use shows how much of that work falls to developers. The agent works in an OpenAI-hosted browser. When it tries to visit a website, the application receives a request it can approve, deny or cancel, and signing in goes through a separate flow. But the guide says approving a site does not force a confirmation before each purchase or destructive action. Applications that need that guarantee must limit what the agent can reach or run a browser they control. Under this design, the consent step Weinstein recommends is something builders have to add themselves.

The agent as its own tester

Asked how computer use is changing software work, Weinstein named one of his favorite use cases: letting the agent test the software it has just built. Normally, he said, Codex writes the code and then "you are now like QA for the agent." QA means quality assurance, the job of checking that software works. With computer use, the agent can build the software and then test it, so "by the time it comes to me, it's already working." His version has an extra layer, because he is sometimes developing computer use itself: he has a computer use agent using his computer use agent, which is using something else. OpenAI's Dots announcement describes a similar workflow. An assistant scopes small fixes from customer feedback, implements and tests them, and presents them for human review with videos of the changes.

The conversation added two more uses. One was a "visual play test" skill, a reusable set of agent instructions, that catches design problems you would not spot by reading the code. Conventional tools such as Playwright check visuals differently: they compare new screenshots against saved reference images, within a set tolerance. The other use was blunter. If you are stuck paying for a poor software-as-a-service product, the agent can drive it screen by screen, take screenshots and notes, and hand everything to Codex to build a clone.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04

From the conversation

Podcast episodes

Latent Space

Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

Episode published This article draws on 2:24–5:01 and 14:34–19:14 (approximate times)

Article history

Updates to this article

Tags