1 October 2026
Heard in AI

OpenAI's Nikunj Handa on API plumbing for agents that run for hours

At DevDay, OpenAI API product lead Nikunj Handa said GPT-6 lets agents keep reasoning while tools run and take new instructions mid-task. He also described longer cache guarantees, one still in preview, and open questions about the platform.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published 30 September 2026

Some AI agents now work for hours at a stretch. They write code, browse, operate apps and wait for slow tools to return. On the Latent Space podcast, recorded on the day of OpenAI's DevDay and published on September 30, 2026, Nikunj Handa explained how OpenAI has changed its API for that kind of work. Handa leads product for OpenAI's API team and says he has been at the company for about three years. The API is the paid interface developers use to build OpenAI's models into their own software.

Handa described his team's job simply. With each new model, it works with the research and post-training teams to find out what the model can newly do, then exposes those abilities to developers. For GPT-6, he said, much of that work was about agents.

Not waiting for the tool

An agent works by calling tools. A tool is any outside program the model asks to run, such as a search, a test suite or a browser action. Until now the model has usually stopped and waited for the result. Handa said tool calls in products like Codex, OpenAI's coding agent, and Dots, its personal-assistant agents, "take so long" that pausing the model no longer makes sense. With async function calling, he said, a model can "kick off a tool call, keep running, keep reasoning, and then check back in."

OpenAI's async tool calling guide shows how this works. A developer marks a function as asynchronous, the model carries on with independent work, and the finished result comes back tagged to the original request. In the guide's example, the model starts a weather lookup and answers an unrelated packing question while it waits. The guide also sets limits. The developer's own application still runs the job, the feature requires GPT-6 Astra or a later model, and it does not apply to OpenAI's own built-in hosted tools.

Changing course mid-task

The second feature is mid-turn steering. A developer can now insert new messages while the model is still reasoning. Handa gave the example of adding instructions once a tool call finishes. The conversation raised the point that this is partly a training problem, since the model has to be taught to accept interruptions well, and Handa agreed. Steering had existed in OpenAI's apps before. The discussion noted that it "wasn't the best" there but has improved a lot.

That timing is deliberate, Handa said. His team's goal is to put a capability into the API only once it has been trained into the harness, the software loop that runs a model through a task. "We kind of wait for that moment until it's good enough," he said. The steering guide states the limits. A steering message changes what the model does next. It does not rewrite output already delivered, undo earlier actions or cancel tools that have already started.

The connection underneath

Both features rely on WebSockets, a kind of connection that stays open so the developer's app and the model can send messages in both directions. Handa said OpenAI launched WebSockets in the API a few months ago. Their first use was GPT-5.3 Codex Spark, a model name he joked he could not believe OpenAI chose. The point, he said, is to "really reduce the overhead of going back and forth with tools." He noted that this is separate from GPT Live, OpenAI's real-time product. OpenAI's WebSocket documentation says later requests send only new input instead of the whole conversation. It reports up to roughly 40% faster end-to-end runs in observed workflows with 20 or more tool calls.

UltraFast, and a question left open

The conversation turned to UltraFast, a premium speed tier that was described as available in the API for the first time. According to OpenAI's guide, developers choose it per request and are advised to use WebSockets, because opening new connections over and over can eat up the speed gained.

Handa said the most fun part had been "watching the inference team cook with Astra." The team kept Codex agents running constantly to squeeze out more performance. For a couple of months, he said, the focus was efficiency and cost, and much of the roughly 80% cut in the price of Luna came from those inference improvements. Then the team "shifted gears" toward making models like Astra run as fast as possible.

A question about hardware followed. Spark had been publicly credited to the chipmaker Cerebras, and the question put to Handa was whether UltraFast is also related to Cerebras, with a note that OpenAI has its own silicon too. The conversation moved on to the Decisions API without an answer. OpenAI's Spark announcement from February 2026 does name Cerebras' Wafer Scale Engine 3 as that model's hardware. OpenAI's UltraFast documentation does not say what hardware the tier uses.

Caching for threads that never end

Asked where he wanted feedback, Handa said the Responses API, which he called OpenAI's workhorse, is now focused on performance. First, the team has been rewriting it to cut latency. Second, it is working on caching. He was thinking especially of personal agents that are "basically, like, a single thread that just goes on and on forever."

Prompt caching lets the model reuse work it has already done on the unchanged start of a long prompt. That makes later requests faster and cheaper. Handa said OpenAI now guarantees cache hits within 30 minutes. OpenAI's caching guide documents a default minimum lifetime of 30 minutes after the cache was last written or used. The guide also warns that rewriting earlier messages breaks reuse. Handa added that OpenAI had launched a 12-hour caching guarantee for one of its users. It is still in preview, he said, and not yet public. Customers pay a little more for the cache, so a thread they return to three or four hours later still gets the caching performance.

His advice to builders was to make their apps "very cache aware" and to use OpenAI's cache diagnostics to find where reuse drops off. That tool compares two requests to show what changed, such as a renamed tool. He also described pre-warming. If a developer knows a prompt is coming, they can pay the cache-writing fee in advance and have the cache ready for the next 30 minutes. The conversation noted that one warmed thread could then be split into many copies, and Handa said you can "have tons and tons of that."

Squeezing the context

Caching does not remove the limit on how much a model can hold in its working memory, called the context window. In the conversation that limit was put at about a million tokens, the word fragments models read. A long-running agent still needs compaction, which means shrinking its earlier history so it can keep going.

Handa said OpenAI has its own proprietary compression. In the Agents API, it is built into the harness. In the Responses API there are two options. With server-side compaction, the developer sets a token threshold and the API shrinks the context automatically once it is crossed. With slash compact, developers get full control and can decide for themselves when to compact. OpenAI's compaction guide describes both.

Handa said a lot of the big coding agents like to compact manually, including Codex in its open-source harness. In Codex, /compact replaces earlier turns with a short summary. He added that OpenAI is experimenting with new compaction techniques, including "file-based systems", and that some are already in the Codex harness.

How much should the platform do?

The Agents API, which OpenAI announced in public beta on September 10, gives developers access to the Codex harness, which OpenAI operates for them. Handa said a meeting-notes feature he discussed is "fully built on top of the Agents API." OpenAI's help page describes its Meetings plugin as producing summaries and action items from captured conversations.

Handa's open question came from his time at Stripe, the payments company. There, he said, much of the work was building higher-level products on top of basic payment building blocks. OpenAI has tried this before. The Assistants API "wasn't really the right fit," he said, and with the Agents API the question is how much flexibility to give developers. He wondered aloud whether OpenAI should offer "memory vaults" and other higher-level objects that hide storage details. Today, memory lives inside products: OpenAI's Dots documentation describes each Dot keeping its own saved notes. Handa said a common pattern in AI is to offer a low-level building block plus an example harness, then let a developer's coding agent build the rest. "How much of that do we build into the API is, like, a constant question," he said, and he asked for ideas.

The conversation then offered one way to think about it: that OpenAI is building an "AI cloud", a framing tied in the discussion to something "Sam" said a year earlier. That means AI-native versions of the building blocks Amazon Web Services created, such as EC2 for computing and S3 for storage. Pre-warming caches for work you know is coming was given as one such new building block. Handa agreed: "Yeah, absolutely."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12

From the conversation

Podcast episodes

Latent Space

Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

Episode published This article draws on 19:14–23:38 and 31:51–39:02 (approximate times)

Article history

Updates to this article

Tags