1 October 2026
Heard in AI

OpenAI built a Jev-style Decisions API in a week, says Nikunj Handa

OpenAI's Nikunj Handa says the Decisions API, inspired by TypeSafe's Jev, began about a week before DevDay and runs on existing Luna weights, not a new model. He agreed confidence calibration may need future model work.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published 30 September 2026

OpenAI's new Decisions API began about a week before the company showed it at DevDay, according to Nikunj Handa, who works on product for OpenAI's API. "This whole thing started, like, a week ago. So, it's, like, very, very early," he said. The team trained no new model for it. Engineers built it on the weights of Luna, a model OpenAI already had.

Handa described the sprint on Latent Space, recorded at OpenAI's DevDay and published on September 30, 2026. Hosts Swyx and Vibhu also spoke with Ari Weinstein, who leads product and engineering for OpenAI's computer-use agents. A large part of the conversation was about one question: if OpenAI could build this so quickly, how close is it to Jev, the product that inspired it?

What a "decision model" is

Most people meet language models as chatbots that write paragraphs. A decision model is meant for a different job. Its answers go straight to other software, not to a person reading prose. TypeSafe's documentation describes its model, Jev, this way. A program sends some shared information, such as a support ticket, plus a list of typed questions. Jev answers each question separately and in parallel. It returns a value and a probability spread across the possible answers, not a conversational reply.

TypeSafe offers three kinds of question. A Choice picks from fixed options, a Score rates something on an ordered scale, and a third type estimates whether a statement is true. The company recommends asking one narrow question at a time and letting the developer's own code combine the answers, as its primitives guide explains.

OpenAI's version follows that pattern. The DevDay recap says developers supply text or image context plus questions with a fixed set of possible answers. OpenAI suggests using it to classify content, route requests and choose an agent's next action. The company announced it in limited preview, with a wider release planned for the coming days.

A week, two engineers and no new model

Handa first gave credit to the competitor. He thanked Diogo and the Jev team for "really inspiring the whole segment." When Jev came out, he said, "everyone's, like, losing their minds over it". OpenAI's API customers were asking for something similar, and OpenAI's own teams wanted "a much faster classification system."

The prototype came from what Handa called OpenAI's "strong, like, hacker culture." Two engineers, one from the Inference team and one from Infra, got "nerd sniped," he said, meaning hooked on a problem too interesting to leave alone. "We're going to, like, hack on it. They build a prototype. It, like, works." At the time of the recording, the team was "hill climbing on latency," making small repeated improvements to speed. He said OpenAI planned to launch the API in the coming days, once it met its latency target.

Handa listed what is actually new in the build. "We're, like, building this purely on top of the same Luna weights that we have," he said. Weights are the learned numbers that make up a trained model. Three changes sit on top of them:

  • Constrained output. The model can only answer in the format allowed by the question. Handa called structured output "a big part of it."
  • A faster inference stack. Inference is the work of running a trained model to get an answer. The team tuned it for time to first token, the delay before the model starts producing output.
  • Parallel questions. A request can contain several questions, and the system runs them together as a batch.

Handa presented this as a first try, not a finished design. The first version, he said, is "zero-shotting this on top of Luna to see how it goes." Zero-shotting means trying something with no special training for the task. He linked the approach to "OpenAI's, like, classic iterative deployment thing. Put it out there. See what people think. And then, like, we'll make more model improvements as needed." The episode's publisher put it more bluntly, describing the product as "just a Luna wrapper for now."

Weinstein had made a similar point earlier in the episode, when the hosts asked whether OpenAI's computer-use agents rely on the Decisions API. Weinstein said the Decisions API runs inference in parallel, does no reasoning and uses a smaller model than the computer-use models. Those choices make it very fast but "a little bit less good at doing like long horizon sort of sophisticated tasks." How to combine the two approaches, Weinstein said, is "still an open area of research."

Is it just structured outputs?

The hardest question in the conversation was whether OpenAI had built anything beyond a feature it already sold. The discussion noted that many Jev clones had appeared in a short time. One line of argument was that copying Jev's interface alone is easy: a Jev-style API is "honestly structured outputs," a feature OpenAI offered first.

Structured Outputs is an existing OpenAI feature. It forces a model's response to match a template the developer supplies, called a JSON schema, including required fields and allowed values. The challenge was simple: take Luna, which the discussion said the decision model is priced the same as, turn off reasoning and require structured output. Is that a Jev? The answer offered was no. A decision model needs more than speed and a fixed format. Getting the format right does not make the answer right, and OpenAI's own documentation warns that schema-conforming responses can still contain substantive mistakes.

One advantage came up quickly: OpenAI's version accepts images and Jev does not. "We get it for free with Luna," Handa said. TypeSafe's System One documentation confirms that Jev takes text, JSON objects and arrays of text, not images, audio or video.

The calibration gap

The biggest difference raised was confidence. The discussion pointed out that if the API still uses the same Luna weights, the confidence and calibration features are probably still to come. A model is well calibrated when its stated confidence matches how often it is actually right. The argument was that RLHF, short for reinforcement learning from human feedback, pushes chat models toward what people want to hear rather than toward an honest measure of how sure they are.

Handa agreed. "These are going to be, like, the key areas where we may have to, like, hill climb with a future model release," he said.

This is what TypeSafe says sets Jev apart. Its documentation describes models trained so that their probabilities track real outcomes. It also notes that calibration is measured across many predictions and does not guarantee any single answer. In Jev, the confidence figure comes from the probability spread it returns. TypeSafe's examples send uncertain cases to a person and set a stricter threshold for a financial transfer than for clicking through a screen.

How does Jev work? Nobody outside knows

The conversation also turned to how Jev produces parallel answers so quickly. The discussion noted that TypeSafe has not explained it and treated the options as speculation. One guess was a diffusion model, which refines a whole block of text at once instead of writing one word after another, as standard autoregressive models do. The other guess borrowed from mechanistic interpretability, the study of a model's internals: reading the model's internal activations directly and turning them into answers. The discussion said demos of both approaches existed, including a Google diffusion demo, and called it "all speculation."

The Google demo was likely DiffusionGemma, an experimental model Google announced in June 2026. It refines blocks of 256 tokens at a time. Google reports up to four times faster generation on dedicated GPUs, but lower overall quality than standard Gemma 4, and says the speed advantage shrinks for busy cloud services handling many requests at once.

Handa did not pick a side. He said he was glad the field had been "kicked off" and that "everyone's going to learn from each other."

What OpenAI is using it for

Handa said the main early use is "really fast classification." Inside OpenAI the first interest was obvious: "the user ops team was, like, jumping on it. We were, like, we were going to classify all of our support tickets." The conversation also suggested using fast decisions as an automated judge for testing AI systems. Handa said that would "be interesting to see."

Computer use, meaning AI operating a computer's interface the way a person does, is a more limited fit. Handa said there will be limits to "having Luna pick, like, one action at a time," compared with having OpenAI's larger Astra model write a JavaScript script to control a computer. "It's not going to be at the same intelligence level," he said, though some computer-use tasks might not need more.

The prototype he seemed most excited about connects the API to GPT-Live, OpenAI's real-time voice system. He described GPT-Live as a "thinker, talker" setup. A fast voice model handles the conversation and hands harder work to a model like Astra in the background. OpenAI's GPT-Live guide describes the same split: the voice model listens and speaks at once while a separate backend handles tools and longer tasks. Tool calls have always felt slow in GPT-Live, Handa said. People inside OpenAI have now put together demos in which GPT-Live controls a computer through fast decision calls, and he said they feel "so much more snappy and natural."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08

From the conversation

Podcast episodes

Latent Space

Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

Episode published This article draws on 1:34–2:24, 23:21–23:37 and 23:38–31:51 (approximate times)

Article history

Updates to this article

Tags