October 2, 2026
Heard in AI

Jev shows language models needn't just write text, Alex Zhang argues

MIT researcher Alex Zhang argues that TypeSafe's Jev, which returns fast typed decisions instead of prose, matters as a test of new language-model designs, not as a mere classifier. He says the open question is how it was trained.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published October 2, 2026

When TypeSafe AI released Jev, a model that answers with a bounded decision instead of free-flowing text, part of the online reaction was dismissive. According to Alex Zhang, the MIT researcher best known for Recursive Language Models (RLMs), critics called it something "we've known for years": just a basic classifier of the kind students build in introductory machine-learning classes.

Zhang thinks that reading misses the point. On the Latent Space podcast, in an episode published October 2, 2026, he argued that Jev raises a larger question: "are language models correct? Like in the form that they're in, can we consider a different design space other than text to text?"

Zhang said he had watched the same cycle with his own RLM work: something gets overhyped, and then people decide it is trivial. Teams have to market their research, he said, and he understands why. But the backlash, in his view, can hide the interesting part.

What Jev does differently

Most of today's chatbots are what researchers call autoregressive decoders. They produce an answer one small chunk of text, or token, at a time, and each new chunk depends on the ones before it. That makes them flexible, but also slow and expensive when the task is narrow.

TypeSafe introduced Jev on September 15, 2026 as a model for typed decisions: a program can ask it for a classification, a choice among options or a rubric score, and get back probabilities it can use to decide what to do next. The company describes parallel probability outputs and reports response times of 70 to 500 milliseconds for these queries. It also says its short-input demonstration favors Jev's design and that its largest advertised speed and cost gains sit toward the high end of what users should expect.

Zhang described the idea this way: keep the language-model backbone, which already captures a lot of information about language, but change its output space, meaning the set of things the model is allowed to produce. If you already know the answer must be yes or no, he asked, why pay for a full text generator? "Am I going to ask my language model to do this and pay like a 400x cost? No, it's like silly, right?" The 400x figure was his own illustration, not a measured result.

Why the default went unquestioned

Zhang's explanation for why people rarely ask this question is about who controls the best models. Because the leading labs make the strongest systems, he said, people use whatever those labs ship, and so they have become accustomed to the idea that a language model is just an autoregressive decoder.

He heard a version of the same objection about RLMs. A frequent criticism, he said, was that "this is not a language model." His reply: "a language model is just modeling language. It doesn't have to be this transformer decoder."

What Jev adds, in his view, is a new dial to turn: what the output space is, and how that choice affects inference latency, the time a model takes to respond. Zhang linked this to looped transformers, models that run some of their layers repeatedly. That idea also drew "this is a silly idea" reactions, he said, but it opens questions he thinks PhD students should want to answer. What if only part of the model loops? What if a router sends inputs to different parts of the model? Could some of what developers now build into a harness, the software loop wrapped around a model, be mimicked inside the model itself? "We don't know how far we can take this," he said of both Jev and looped transformers.

Speed for swarms

Zhang said his first reaction to Jev was that it would be "really useful for RLMs." In RLMs, swarms of agents and similar systems, a model calls other models over and over, and the biggest bottleneck, he said, is that they are slow. Some of those calls are trivial, but they still go to a large, "bulky" model because that is the only kind available. A cheap, fast model for the simple steps would let such systems spread their computing more sensibly.

He predicted that new types of models will emerge "beyond just the bog standard frontier model." The conversation also brought up the interaction model from Thinking Machines, nicknamed "Thinky," and Zhang called it "another great example." Thinking Machines' preview of that work describes a model that takes in any subset of text, audio and video and predicts text and audio, working in continuous 200-millisecond micro-turns, while a separate background model handles longer reasoning.

Why newer labs have to bet differently

The discussion turned to competition. One point raised was that a company trying to beat a frontier lab at building autoregressive decoders would be outmatched on computing resources. Zhang said a newer lab that simply replicates OpenAI or Anthropic is following "a horrible strategy," because it ends up competing on data and compute, the two things it cannot match. He added that he did not know much about Thinking Machines and was not saying that was its plan. (Thinking Machines has its own large model, Inkling, an open-weights release from July 2026.)

"Anything other than OpenAI, Anthropic, maybe like Meta and GDM, like you just, you got to do something else," Zhang said, calling it "the sad reality." GDM is Google DeepMind. But he also called it a good thing: because scaling works, the big labs will keep doing it, leaving room for new players. He doubted that frontier labs are seriously pursuing alternative designs: "why would you take the risk of allocating a large amount of compute to new bets when the old bet already works?" When the conversation suggested that this is what spins off neolabs, researchers with side bets who get no compute and leave, Zhang answered, "Exactly."

The hard part is training

Zhang said he had some guesses about how Jev works and had seen people say it is a diffusion model; the discussion linked that to parallel decoding, which produces many outputs at once. But he said what excites him is the training, not the architecture. "I'm not entirely sure what their optimization objective was and how they trained it," he said.

He compared this to RLMs. The idea is simple enough that anyone can use it once the paper is out. The real value, he said, is whether you can train the system properly and maybe mold an architecture around it. He said TypeSafe had "figured out a way to train the system, which is completely non-trivial." He had seen claims online that Jev was post-trained on top of Alibaba's Qwen models, and said any open-source Qwen setup people try will be worse, because whatever TypeSafe did "clearly works very well." TypeSafe's announcement names its method Reinforcement Learning for Calibrated Decisions.

Calibration

The conversation then turned to calibration: whether a model's stated confidence matches how often it is actually right. One point raised was that if you ask an ordinary chatbot how confident it is, it produces a likely-sounding answer such as "43%" rather than a number tied to its real accuracy, whereas Jev gives a grounded classification. The discussion also noted that many people use it only as a fast classifier without using its probability estimates, and suggested calibration is fairly easy to build synthetic training data for: take questions with known answers, have the model classify them and compare its outputs with the truth.

TypeSafe's announcement makes the same distinction, saying confidence should correspond to observed accuracy. Its published workflow evaluations, though, measure Jev against averaged probabilities from two other models, Astra and Fable 5.1, not against ground-truth labels.

Games as a test

Zhang said Jev will not solve everything, but it addresses a class of problems AI has long struggled with: low-latency tasks. He said he loved TypeSafe's game examples, and the conversation singled out its Doom demo. In TypeSafe's JevDoom demo, the model plays an original arena game, not Doom itself. It receives structured information such as enemy direction, ammunition and nearby walls instead of screenshots, chooses among ten actions, and the game pauses while it waits for each decision.

Games matter to Zhang because he built a benchmark for them. VideoGameBench, released with colleagues in May 2025, tests whether vision-language models can play well-known games using only screenshots, basic objectives and controls, with a minimal harness and the game's real-time pace included. Zhang said it came out soon after Claude Plays Pokémon. In the original paper's real-time tests, the best overall score, Gemini 2.5 Pro's, was 0.48% average progress, all of it from the Kirby game. Zhang said those numbers are now "very outdated" and that others are running newer models on the games.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

From the conversation

Podcast episodes

Latent Space

Academia is for Ambition — Alex Zhang, MIT

Episode published This article draws on 19:20–29:45 (approximate times)

Article history

Updates to this article

Tags