October 9, 2026
Heard in AI

Frontier-lab leader called a compute cap 'reasonable,' Labenz says

Nathan Labenz says an unnamed frontier-lab leader told him there is likely a level of intelligence we shouldn't pass, and called a limit on the next pre-training run's compute "reasonable." Labenz himself adds that counting compute is hard.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 8, 2026

A leader at one of the frontier AI labs said there is probably a level of intelligence we shouldn't go past. That is the account of Nathan Labenz, host of The Cognitive Revolution. When Labenz pushed for specifics, he said, the executive called a hard limit on the computing power behind the next big training run something that "could be reasonable."

Labenz heard it at The Curve, an invitation-only conference held October 2–4, 2026. Its organizers describe a gathering of researchers, policymakers, executives and people with differing views on AI acceleration and safety, and it includes a Chatham House track, where participants can repeat what was said but not who said it. Labenz said his conversations with founders, executives and top researchers at frontier companies were all under those rules. He did not name the executive, saying only that this was a person whose "name and position" everybody would recognize. What follows is Labenz's retelling, given on the podcast alongside his co-host Prakash Narayanan.

"A level of intelligence that we just shouldn't go past"

Labenz called the remark "legitimately newsworthy." In his words, the executive believes "there likely is, let's say, a level of intelligence that we just shouldn't go past."

The reaction in the conversation was a single word: "What?"

Labenz said it was the first time he had heard anything like it from a frontier-lab leader. The executive did not claim to know exactly where that level is, and the statement "was not like minced words." Labenz added that it was "a little bit ambiguous" whether the executive meant never or just not right now. Still, he called it "definitely the strongest statement I've heard about imposing kind of a hard cap on capabilities, even for a time."

The objection: capping intelligence means capping its inputs

Labenz introduced Narayanan's reaction as being about what a ceiling would actually require. Recent gains in AI, the argument ran, come from two sources: more computing power, which has driven progress for about a decade and a half, and better training data, which is increasingly refined by more capable models. "When we say that we shouldn't go exceed a point of intelligence, that automatically means that you have to look at the inputs going in," the argument went. A ceiling would mean not building compute past a certain point, or not refining data further. The conclusion: "It means the end of the compute build-out, which is pretty significant."

Testing it with a number: 10^27 FLOPs

Labenz said he followed up with the executive. Would he sign on to a limit on the number of FLOPs going into the next pre-training run? A FLOP, short for floating-point operation, is one basic arithmetic step a computer performs, and the total count is a standard way to measure how much computing went into training a model. Pre-training is the first and most compute-hungry stage of building a model. Labenz floated something like "no more than 10 to the 27 flops in the next pre-train," while saying he did not know what the right number would be.

He expected the executive to treat the figure as a straw man, explain what was bad about it and ideally offer something better. Instead, Labenz said, "he basically said, yeah, I think that could be reasonable. The end, not really much of a fight there at all."

The catch, in Labenz's own commentary, is deciding what counts. Labs spend computing power generating synthetic data, meaning training material produced by models. Should those operations count toward the cap? "Maybe you should," Labenz said, calling the question of how to count "super vexing." The more compute goes into enriching data, he warned, the more companies could "play a sort of shell game of hide the compute."

The reply: compute can go into safety

Labenz's answer was that a cap on pre-training, or another compute cap acting as a "pacing mechanism," would not end the build-out. "I would say no," he said.

His argument drew on a point he attributed to Jensen Huang, NVIDIA's chief executive, in an interview with Ezra Klein. As Labenz recounted it, Huang said NVIDIA spends about 20% of its effort designing a chip and 80% verifying, validating and testing edge cases so it is reliable and lasts. In Labenz's reading, Huang's message to AI companies was that they had succeeded in making models capable enough to be useful, and the next era would be about making them safe, reliable and trustworthy.

Labenz said he asked one of the lab insiders about Huang's take. The response, he said, was basically that it was "kind of reasonable": without knowing where the numbers would land, the majority of future compute could go into safety measures. Labenz listed several kinds:

  • Monitoring model reasoning. Chain-of-thought monitoring means reading the step-by-step reasoning a model writes out before answering. Labenz said people are getting "kind of bearish" on it.
  • Monitoring model internals. Activation monitoring looks at the model's internal signals rather than its written words. Labenz said labs already spend a significant percentage of compute on monitoring and that it sounds like it "might go up a lot."
  • Fixing training environments. Labenz said labs already spend compute on this in a big way.

Paying models to break their own training grounds

The third item needs background. In reinforcement learning (RL), a model practices tasks inside a training "environment" and is rewarded when it succeeds. If an environment is sloppy, a model can collect the reward by cheating, for example by exploiting a loophole instead of solving the problem. This is known as reward hacking. "If you have sloppy RL environments that reward cheating, then you're going to get a lot of cheating," Labenz said.

According to Labenz, labs now point their models at the environments themselves with a direct task: "hack this environment." The models find the holes and the labs fix them; Labenz said he was sure they are also throwing out environments that are fundamentally flawed. Fewer flaws mean fewer chances for cheating to be rewarded, and so less cheating from the finished models.

Labenz said he got the sense that labs see "a pretty clear relationship" between cleaner environments and lower cheating rates downstream, though he called it "an open research question." It is not yet a scaling law, he said, but "a proto scaling law": a lot of compute spent driving flaws down gets better behavior at the other end. Zero hackable environments would be the goal, he said, but "that's going to be tough."

Bottlenecked on safety

Labenz said he "heard repeatedly" that the labs expect to be bottlenecked on safety and alignment, the work of making models behave as intended. For at least the next generation or two, he said, there still seems to be enough low-hanging fruit in that area, partly from fixing RL environments, that labs should be able to get comfortable releasing the next couple of models. Past that, he said, the answer he got was: "all bets are off."

Share this article

Go to the original

Sources & further reading

  1. 01

Connected ideas and articles

From the conversation

Podcast episodes

The Cognitive Revolution

AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Episode published This article draws on 0:00–0:56, 4:13–5:42, 7:57–11:26 and 18:55–23:09 (approximate times)

Article history

Updates to this article

Tags