Noam Brown, a researcher at OpenAI, described a scheduling problem he expects AI safety testing to run into soon. Suppose an AI model can work effectively on a task for three months. If a new frontier model comes out every two months, the lab never gets to watch the current one through a full three-month job before its successor replaces it.
Brown raised the scenario on the Dwarkesh Podcast, in an episode published on September 17, 2026. He treated it as a problem that is coming, not one that exists now. "This isn't an issue right now," he said, "but it is quickly becoming an issue that we have to figure out a solution for."
The obvious fix is to slow down releases and allow more time for testing. The conversation then turned to the cost of that fix. The longer a lab holds a model back, the bigger the gap between what the lab can use and what everyone else can use. Brown agreed, and pointed to his own company's unreleased mathematics model as an example.
Why longer tasks break the testing schedule
The discussion began with recursive self-improvement, or RSI: AI systems doing the research that produces better AI systems. The question was how anyone would know, at each step of that process, that the models were still aligned, meaning they still behave as their developers intend and serve human interests.
Brown's answer started with speed. He said frontier models are now released "at most every two months," and that a new AI breakthrough sometimes arrives every week. He added that people who last looked closely at AI six months or a year ago would find today's models "far beyond" what was possible then.
At the same time, he said, AI agents can work on their own for longer and longer stretches. Agents are AI systems that plan and carry out multi-step tasks. Before any release, labs run safety evaluations, tests meant to confirm that a model is properly aligned and "in good shape." Brown said labs have worked this way since around GPT-4. The practice quietly assumed that such evaluations could be done in a fairly short period.
That assumption weakens as tasks get longer. According to Brown, today's models can already handle a week-long task, and "we'll probably get to the point where they can do month-long tasks," and later three-month ones. Once a model can work longer than the release cycle, he said, "you don't have a way to evaluate the models at the full length of their capabilities" before the next model arrives.
He did not say which problems would show up over those long stretches, only that nobody would have had time to look for them. The model's capabilities might degrade over a long run. So might its alignment or its safety behavior. Brown said this is not only a safety question but also a product question, since the product itself might get worse over a timespan that has never been properly tested.
He also said many safety policies were written in the GPT-4 era, when this problem was "not on anybody's radar," and that for a lot of companies they have not really been updated since to account for agents that work over very long horizons. In his view, not enough people, inside or outside the labs, are thinking about how to prepare. "If you just look at the trend lines," he said, "we're going to hit this."
The other side: models that stay inside the lab
The conversation then argued the other side. During RSI, work that now takes three months might take one. If so, a delay that looks short on the calendar could hide a large difference in capability between a lab's internal models and its public ones. And if a lab gets enough value from using its models for its own research, it might ask why it should build classifiers and other safeguards, and take criticism, just to release a model publicly. It might stop releasing models at all during RSI, reasoning that there is no point in helping rivals do RSI with its own models.
The result described was a "tremendous concentration of power." The discussion noted that the outside world already lacks access to the models behind the latest mathematical results. It argued that such models will eventually matter far beyond mathematics: to political leaders making decisions, to the media, and to people running businesses. By default, the argument went, public releases will lag well behind what labs use internally once progress speeds up significantly.
Brown called that "absolutely right." He laid out how a reasonable safety argument leads there. Models are becoming extremely powerful and are working over longer horizons. Labs want enough time to evaluate them before release, so the release cycle should slow down. "And there's a flip side to that," he said: slower releases create "more of a disparity between what is internal to the labs" and what the outside world can use.
Mathematics as the first case
Brown said mathematics is the first field where this problem is clearly visible. He described "a very powerful model internally that is currently not available to the outside world" that can solve "incredible math problems." He said its results are not limited to Millennium Prize problems, a set of famous open problems in mathematics; people have also used it to get solutions to many other unsolved problems.
The Millennium Prize problem in question is the subject of OpenAI's September 8 announcement, which reports a result on the Navier–Stokes problem, a long-open question about the equations that describe how fluids move. OpenAI said the result came from about 10,000 agents working in parallel and was then checked in Lean, a formal proof-checking system. The announcement says the system behind the result was an internal model significantly more capable than GPT-6 Astra. OpenAI said it had been training that new internal model since August 28, and that its training was still ongoing at the time of the announcement.
Brown did not suggest a way out. "We don't have a good answer," he said, calling the situation "an unfair advantage." He said there are trade-offs, that he does not know how to weigh them appropriately, and that there is complexity on both sides.