October 9, 2026
Heard in AI

Can AI go superhuman beyond checkable tasks? Two views differ

Nathan Labenz argues an AI-assembled music track shows models generalizing beyond tasks with checkable answers. Mercor's Edward Hu is more conditional: progress depends on how clearly and cheaply 'better' can be measured.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 8, 2026

A common reassurance about AI's limits goes like this: models become superhuman only where their answers can be checked automatically, such as math with a known result or code that passes its tests. Nathan Labenz, host of The Cognitive Revolution, argued in the podcast's October 8, 2026 weekly highlights edition that this wall is not going to hold. His evidence ranged from what he heard at The Curve conference to a music track made with an AI model that cannot hear. Edward Hu, who leads AI modeling at Mercor, offered a more conditional answer.

Why checkable answers matter

Much recent progress comes from reinforcement learning, or RL: a model attempts a task, receives a score, and is trained to do more of what scored well. That works most cleanly when the score is cheap to compute and hard to argue with. A proof is valid or it isn't; a program runs faster or it doesn't. Tasks like these are often called verifiable. The open question is what happens in fields where nobody can write down a reliable score.

What Labenz heard at The Curve

Labenz reported off-the-record conversations with people at frontier AI companies at The Curve. One sentiment he heard framed the question as whether AI will be superhuman "at just some things like the verifiable tasks" or at everything. The middle position, which Labenz called "very credible," was that the labs are seeing superhuman performance at anything they care about and are really committed to investing in.

That is not full generalization to every domain, he said. But by his account, even in areas not thought of as inherently verifiable, the labs believe a familiar package works when they focus on it: license whatever data they need, process it heavily to augment it and create synthetic versions, and apply RL on top. That package, Labenz said, is broadly understood to work on essentially any problem they choose to focus on.

The next target, according to what he heard, is science. One person went as far as to say that in no more than a year, it will not be possible to do the very best science without AI playing a big role in it.

Hu: it depends on defining 'better'

The next day, Labenz put that view to Hu, adding the claim that in some domains human data no longer helps. Hu said such domains exist. His example was kernel optimization: writing the small, low-level programs that make a computation run as fast as possible on particular hardware. There, a model can write code that an expert human might not even understand, and the result is easy to judge because the code either runs faster or it doesn't. As long as the setup cannot be gamed, and Hu said hardware is often very hard to hack, models can reach superhuman performance there, if they have not already.

Where human judgment still matters, he said, things look different. Human taste is "really hard to encapsulate in a number," and in those areas the bottleneck to improvement is still human.

Hu said it comes down to how clearly a field can define what it means to be better, and whether that definition can be turned into a function that is cheap to evaluate and uncontestable. When it can, putting in a lot of compute lets even existing algorithms find good solutions efficiently. But he added that a lot of the economy is not yet there when it comes to specifying what better means.

A track made by a model that can't hear

Labenz closed the episode with what he called something lighter: a track that Prakash Narayanan, Labenz's co-host, made with Claude Fable, built from clips of voices listeners would recognize. Prakash, as Labenz introduced him, explained how it came together. With tokens left over from Fable, he had the model pick seven or eight clips of key moments from the past couple of years. He then gave it access to a digital audio workstation that it controlled through an API. A digital audio workstation is the software musicians use to combine instruments on a track; an API is a way for one program to operate another. The model was told to play the instruments and decide which to use, though it cannot hear.

Prakash asked for a narrative and listened to about nine versions, objecting to parts each time. The model chose the samples, the specific phrases and the sequencing. The result opens with a clip of Dario saying the models just want to learn, which becomes the refrain, and ends on a sober note from Kurzweil about progress toward a human-machine civilization. Each time a voice says "we've got to slow down," the track speeds up, which Prakash described as his own small contribution.

Prakash also drew a line between this and Suno, calling Suno real generative AI and his project more like playing instruments. Suno describes its generator as turning a description of style, mood or subject into a complete song or instrumental track. Prakash's model instead assembled a piece step by step, as a producer would, by working the tools.

Labenz's reading, and the tension

Labenz granted that "on one level, who cares?" The track isn't moving GDP, he said, and could be seen as people "amusing ourselves to death in a new form," a perspective he called somewhat valid. But he said it goes to the heart of a question everyone has been asking: how good these systems will get in domains that are not automatically or readily verifiable. "It really does seem like the generalization is pretty strong," he said.

His argument rests partly on a guess. Labenz said he "highly doubt[s]" the labs run much of a feedback loop on music, and suspects they "just kind of threw some stuff in" and shipped once it was better than last time. If a model can do this well in a domain nobody trained it hard on, he reasoned, its skill is spreading beyond the tasks it was scored on.

That sits alongside Hu's more conditional account. Hu did not say models fail where taste matters; he said progress there is still limited by humans because "better" is hard to put in a number. Labenz's example touches that point: Prakash judged the versions by ear over about nine rounds. For Labenz, a capable result in an unmeasured domain is what matters.

"If you think that was going to be the last wall to hold," Labenz said of the verifiable-reward limit, "I think unfortunately, it's not going to." He then played the track once more.

Share this article

Go to the original

Sources & further reading

  1. 01

Connected ideas and articles

From the conversation

Podcast episodes

The Cognitive Revolution

AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Episode published This article draws on 15:57–16:48, 24:59–27:08 and 1:15:47–1:20:54 (approximate times)

Article history

Updates to this article

Tags