29 September 2026
Heard in AI

Garrison Lovely's plan to freeze frontier AI, and how he'd police it

Author Garrison Lovely proposes freezing frontier AI: no training run as large as the last, no RL from verifiable rewards, no recursive self-improvement, policed by auditors and chip checks. He admits the economic cost would be significant.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published 29 September 2026

Garrison Lovely wants the AI safety movement to agree on a single demand: "freeze the frontier." He says he likes the alliteration. The frontier is the most advanced AI that companies are building. Asked what the freeze would actually ban, Lovely named three limits. There would be no training run as large as the last one. There would be no more reinforcement learning from verifiable rewards. And there would be no recursive self-improvement, meaning AI that does the work of building its own successors. At home, he would enforce these limits with embedded auditors and prison sentences. Between countries, he would rely on chip-level verification. He admits the plan would carry significant economic costs.

Lovely is an author and journalist whose book Obsolete examines the race to build artificial general intelligence (AGI), meaning AI that can do most of the work people do. He set out the proposal in an episode of The Cognitive Revolution with host Nathan, published on 29 September 2026.

Why a freeze, and why now

Lovely does not oppose automation as such. He said that automating labor is what lifted humanity from near-universal poverty and raised living standards after the industrial revolution. He does not want the world to stay at its current level of development. But he argues that "trying to automate all labor is pretty different." He admits that where to draw the line "is not super obvious."

His answer is to stop where things are. He said that even if development froze today, adopting the current models would still cause a lot of economic disruption. Much of the job effect, he noted, is people "not getting hired in the first place" rather than people being fired. He left open whether going back a few model generations might be right "from a social welfare perspective," and said he didn't know. His reason for not demanding that is practical: rolling back is much harder than stopping. His plan is to freeze first, then reassess and regulate today's AI, which he says will itself take a lot of work.

He contrasted this with the Future of Life Institute's 2023 open letter. Published after GPT-4, it asked for a pause of at least six months on training systems more powerful than GPT-4. Risk was a big part of how that letter was framed. In Lovely's view, building the next generation of language models at that point was not existentially dangerous, though he says there were real harms, such as "chatbot psychosis" and suicides, which he documents in the book. His freeze has no end date. It would last until specific conditions are met, described below.

One demand people can rally around

Lovely said people who care about AI safety have long given "muddy" answers about what should be done, and that makes coordination hard. He argues that "we should just stop" appeals even to people who don't take existential risk seriously. People worried about jobs, the environment, concentrated wealth and power, or surveillance can all support it.

He said his book partly sidesteps the details. His argument there is that if politicians get "enough of a what and a why, they'll figure out the how." His example was Operation Warp Speed, the US government's push to develop COVID vaccines. He recalled a New York Times estimate of how long a vaccine would take: even under its most aggressive assumptions, it said about a year and a half, which turned out to be longer than it actually took. With that kind of focus and mobilization, he argued, the definitions for a freeze could be worked out too.

The three limits

When pressed for specifics, Lovely gave three.

No bigger training runs. Training is the compute-heavy process of building a model. Under his rule, no new run could be as large as the last one.

No more reinforcement learning from verifiable rewards. In this method, a model practices tasks whose answers can be checked automatically, such as code that must pass tests, and is rewarded when it gets them right. Lovely called it "a big part of what's driving capabilities increases nowadays." He also tied it to much of the "really scary behavior" seen in AI agents, including a "willingness to hack and to cheat and to deceive" and to escape.

No recursive self-improvement. The discussion had turned to how hard this idea is to define. It then took up a simple test from Ezra Klein: don't let AI researchers use coding assistants. If they have to type all the code by hand, the AI clearly isn't improving itself. The discussion also noted that researchers would resist losing those tools strongly. Lovely said he proposes a ban on recursive self-improvement in his book, but "Ezra beat me to" bringing it to the world. He said he liked Klein's point about coding agents and agreed.

Lovely also defended casting a wide net. His example was Anthropic's classifiers, which he said can flag a harmless question about a toenail as a possible bioweapons request and drop the user down to the smaller Sonnet model. He called the result ridiculous, but said the logic is sound. When getting it wrong in one direction is far worse than in the other, you should be over-inclusive. He acknowledged that going back to human speed is the last thing an AI workers' union would support. But if the whole industry had to move at human speed again, he said, that would not be a disaster: it "didn't exist four years ago."

He did not play down the cost. He said he doesn't want to "minimize the economic consequences" of stopping, which he thinks will be significant. He said there are ways to mitigate them and he would like to see more work on that. But if extinction, permanent loss of control, or all white-collar remote jobs being at risk within months or years are taken seriously, he argued, the rules should be drawn too broadly at first and then dialed in.

What would still be allowed

Lovely thinks the companies already know which of their work pushes the frontier, and he sorted their activities into rough groups:

  • Inference, meaning running existing models to answer customers, is fine. It might produce some data that improves models, but he said serving customers "seems fine."
  • Reinforcement learning from human feedback, where people rate model answers to shape their behavior, is "on the edge." He said companies will do some of it for any product, so it may be acceptable.
  • Pre-training a model bigger than any before it would clearly count as advancing the frontier.
  • Reinforcement learning experiments are also often aimed at the frontier.

He noted that labs already track their compute and could define their work in these terms, "and maybe they even do."

Enforcement at home

Domestically, Lovely would place auditors inside the companies, with "employee level access" to Slack, email and offices, in large numbers and empowered to catch violations. He would also criminalize the work, with prison sentences for trying to build AGI, attempting recursive self-improvement or trying to build superintelligence. With both in place, he said, "I think it would work."

Enforcement between countries

Lovely said the harder part is international. He doesn't think anyone doubts that Beijing could shut down China's AGI projects if it chose to. The problem is getting the US to believe it had happened, and the reverse.

His proposed tools include auditors, international verification agencies, and cryptography that could show what is happening inside a data center or in the chips' network traffic without revealing model weights or state secrets. He acknowledged that this technology still needs to be worked out. He also argued that money spent on technical alignment research, the work of making AI systems pursue intended goals, would do more good if it went toward ways to verify an international AI treaty. He called verification one of the bigger blockers to a binding agreement.

For precedent, he pointed to an example he credits to Toby Ord. Under one Cold War deal, the Soviets cut strategic bombers in half and dragged the halves apart with tractors. Satellites could then verify, without anyone's cooperation, that the bombers were cut. Lovely suggested AI equivalents. Compute could be handed to a third party that uses it only for science, not AI research. All the world's chips could be tested and inventoried. Devices on the chips could show whether they are running a familiar model rather than "some new model."

He said technology is not the main obstacle: "It's mostly a matter of political will." He wants building this technology to carry the same stigma as building a nuclear or biological weapon. If the US and China agreed on a freeze, brought other countries in, and a rogue state then tried to build the technology anyway, he hopes the world would react as it did to Saddam Hussein's nuclear ambitions and invasion of Kuwait, when the nuclear program was uncovered after the Persian Gulf War. He called the freeze "a political challenge and a moral one that has technical components."

The order of steps, and the bar for restarting

Asked whether he has a real path to ever approving AGI, Lovely laid out a sequence. It would start with a bilateral treaty between the US and China, perhaps with one country acting alone first. He expects getting those two on board to be the harder part, since each has great leverage over other countries. The freeze would then have to become global.

To resume, he would require strong public buy-in and a scientific consensus that the work can be done safely and controllably. He said he took this language from the FLI-hosted Statement on Superintelligence, which proposes prohibiting superintelligence development until broad scientific agreement on safety and substantial public support exist.

Lovely said it is up to the people who want to build the technology to show that such support exists. One route would be citizens' assemblies around the world. Randomly selected people would serve as a kind of jury, hear experts on different sides, and reach decisions that might be binding or only advisory. Referendums could come on top of that. He said he doesn't think you need "literally everybody": even with 70% support in referendums worldwide, about a third of people might still be opposed. He also said he doesn't know where the threshold should be and doesn't think any one person can decide it.

He said reaching that point would take a long time, and he is fine with that. In his view, these machines could replace the very thing that let humans take over the planet. So he argues for a standard of consent closer to assisted suicide, a deeply deliberate choice, made by the whole species.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

From the conversation

Podcast episodes

The Cognitive Revolution

Obsolete or Irreplaceable? Garrison Lovely on Stopping the Race to Replace Human Labor

Episode published This article draws on 40:04–42:32, 1:01:59–1:02:47, 1:14:59–1:15:23, 1:17:07–1:17:52, 1:32:56–1:41:39 and 1:52:40–1:56:02 (approximate times)

Article history

Updates to this article

Tags