25 September 2026
Heard In AI

Robinhood's Vlad Tenev on why AI research may be automated before consumer apps

Anthropic reports that in August, 26% of the weighted work in a fixed sample of its research and engineering tasks had reached a level where Claude does most of a task under human supervision; no category was rated fully autonomous. On Moonshots, computer scientist Alexander Wissner-Gross projected from that trend that AI could be leading all of its own research within three to twelve months. That is his forecast, not a conclusion of Anthropic's study. Robinhood CEO Vlad Tenev argued that AI research is easier to automate than consumer products because its experiments come with ready-made tests, while app changes must wait for real users to produce statistically significant results.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published 19 September 2026

In an episode of Moonshots with Peter Diamandis published September 19, 2026, host Peter Diamandis, founder of XPRIZE and Singularity University, turned to what he called "the recursive self-improvement story of the week." Recursive self-improvement is the idea that AI systems help do the research, software and hardware work that produces the next, better AI system, which could then speed up the work again. His news item was a disclosure from Anthropic: its Claude models, he said, now lead roughly 26% of the company's measured AI research and development work.

The panel read that number in two ways. Alexander Wissner-Gross, a computer scientist and founder of Reified, treated it as a trend line pointing toward AI running its own research soon. Robinhood co-founder and CEO Vlad Tenev asked a different question: why would AI research be among the first kinds of work to be automated, ahead of the apps people use every day?

What Anthropic actually measured

Anthropic's methodology page makes the 26% figure narrower than it sounds. The company built a fixed "basket" of research and engineering work. In each week of July, researchers sampled 20% of staff in the relevant departments. Claude then went through their work records and identified roughly 15,000 tasks, sorted into 378 categories. Claude-based judges gave each category an automation level, and the results were weighted by how much staff time each kind of task was estimated to take.

The 26% refers to level four, or AL4, on that scale. At AL4, Claude can do most of a task from a high-level request while a human supervises. For August, 26% of the weighted work was rated AL4. More than 90% of the work reached at least AL3, which means Claude did substantial work under close human direction. No category in the basket reached AL5, the fully autonomous level. Anthropic calls the ratings approximate and notes that a fixed basket can miss new kinds of work as they appear.

The other number Diamandis cited, about 30,000 agents working at the same time, is a separate measure. Agents here are AI systems that carry out multi-step tasks with tools. Anthropic reports that about 30,000 research and engineering agents run at once on its most-used internal platform, and that all their actions are subject to online monitoring. That count describes the platform. It is not what the 26% is a share of.

The speakers also gave different starting points for the trend. Diamandis said the share was up from 1% at the start of the year. Wissner-Gross, describing the chart shown on screen, said it had risen from 3% in April.

Wissner-Gross's three-to-twelve-month extrapolation

Wissner-Gross called the chart "incredible" and said many people had suspected something like it. He expects the automation levels to follow sigmoid curves: S-shaped curves that rise slowly, then steeply, then level off. Extending the AL4 trend along such a curve, he said, suggests that "approximately in the next three to 12 months, depending on uncertainty, AI is just completely leading all of its own R and D," which he called "total recursive self-improvement."

He does not expect a sudden jump, but a curve. His "primary take home," he said, was that "we're already substantially all of the way to recursive self-improvement," at least within Anthropic.

That forecast is his own projection from the data. The study itself describes supervised work: AL4 still has a human supervising, and Anthropic rated no category autonomous. Wissner-Gross also argued that the trend goes beyond one company. He pointed to what he called "a lot of smoke" from Google DeepMind, including a recently released paper on its own self-improvement research that he did not name, and to hints that the next Gemini model will lean heavily on self-improvement to catch up. He said OpenAI has "made no bones" about using it in its releases. Soon, he predicted, people will be asking "what is the role of human researchers anymore in driving new releases?"

Tenev: AI research has tidy edges

When Diamandis asked whether he was seeing the same thing at Robinhood, Tenev answered, "Oh yeah, absolutely." He said recursive self-improvement is coming to every software project and probably to hardware too, though hardware will take a little longer.

He then argued against a common assumption. Many people expected AI research to be automated last because it is complicated. Tenev thinks it is "probably one of the easiest things to automate." In his account, models are sandboxed, meaning they run in a contained environment, although he added that "we can debate whether they're actually sandboxed." Their interfaces are straightforward and they "don't depend on a lot of other things." Researchers already run experiments and score them with evals, automated tests that measure whether a model got better. Automating both the experiments and the evals is therefore fairly straightforward, and the "surface area," as he put it, is "pretty contained."

A product like Robinhood has many more moving parts: the iOS app, the backend systems and, at the end of the day, the humans it is built for. Nobody has yet found a good way, he said, to replicate how people will respond to a new feature, or to know in advance whether a metric will improve by a statistically significant amount, meaning with enough evidence that the improvement is probably not due to chance.

He still expects consumer products to be automated end to end eventually. The bottleneck, in his view, will be "how quickly can you get statistical significance that a change is an improvement over the status quo," so that the change can go into the live product instead of being discarded. The answer depends on how many people use the product. That, Tenev said, favors large platforms, "somewhat sadly." A company like Meta, with billions of users, can tell very quickly whether a change works. A company with fewer users has to wait longer for the same certainty. In Tenev's account, the wait comes from people, not from how quickly changes can be written.

Blundin's illustrations: dense math, small code

Dave Blundin, founder and general partner of Link Ventures, had given an example earlier in the episode of how a single mathematical idea can make AI itself run better. It came from studying Moonshot AI's Kimi K3 model on a flight back from California the day before. Large language models rely on attention, the mechanism that lets the model weigh earlier words when producing the next one. To avoid recalculating everything at each step, models keep stored data about earlier words, called the KV cache, and that memory takes up expensive hardware. Blundin said Kimi's KDA attention method, short for Kimi Delta Attention, found a way to cut out "three quarters" of that cache.

His reconstruction of how the researchers got there is his own reading of the work. In his telling, they ran a thought experiment: what if they skipped softmax, a normalizing step normally applied to part of the attention calculation? They worked through how the math changed, simplified it heavily, then ran a quick test to see whether the cheaper approximation performed as well as the harder-to-compute original. Blundin thinks an AI agent could now do all of that "just thinking through the math." For him, this is why mathematics ties directly into AI optimizing its own performance: a simple mathematical adjustment can ripple all the way down to Nvidia GPU performance "at the transistor level," which he called "exactly the same process" as the self-improvement loop.

Moonshot's Kimi K3 model card shows that KDA is one part of a mixed design, not a full replacement. Of the model's 93 attention layers, 69 use Kimi Delta Attention and 24 use a different mechanism, gated Multi-head Latent Attention. Moonshot presents KDA alongside other architectural changes meant to support long inputs and agent-style work.

When Tenev laid out his bottleneck argument, Blundin tried to put numbers on it. He said he had re-implemented Kimi K3, and a team member had re-implemented GLM, another AI model, in order to speed them up; that came to about 10,000 lines of code. He guessed that Robinhood might run to 30 million lines. Tenev wasn't sure: "I don't know if it's quite that much." Blundin replied that a core portfolio accounting platform alone is usually 10 or 20 million lines. By his estimate, "the scale of an AI algorithm is microscopic compared to a major consumer application." He described that code as very dense but "incredibly sandboxed," and said the Kimi K3 code has "literally no loops." It is far easier, he said, for an AI researcher to work on AI algorithms than on Robinhood's, and so self-improvement is "definitely pointing inside of itself first."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Robinhood's Vlad Tenev on Tokenizing Everything, OpenAI's 6 Misalignment Reports, Figure's Robot Makes Beds | EP #292

Episode published This article draws on 48:07–49:34 and 1:50:01–1:56:27 (approximate times)

Article history

Updates to this article

Tags

Box's two rules for software in the agent era: beat the generic agent, then let it in

On Sequoia's Training Data podcast, Box CEO Aaron Levie said any company sitting on customers' data now has two obligations: build an agent measurably better than an off-the-shelf one at its own workflows, and expose the same capabilities to outside assistants like Claude and ChatGPT. He described the tuned search-and-retrieval harness behind Box's agent, the evaluations that track model progress, and his bet that within five years roughly 90% of enterprise tokens will be spent on work nobody asked for directly.

· Updated 8 min read

Huang says AGI has arrived; OpenAI's 3.1 figure answers a narrower question

Nvidia's chief executive declared AGI achieved on September 6 while announcing more GPU capacity, and the Moonshots panel split over whether the label means anything at all. Later in the segment, after OpenAI's Codex lead said the model had moved some plans six months earlier, one panelist called that the most important moment in human history. A second claim discussed on the same show, that OpenAI's agents now do 3.1 days of research work per human day, comes from an internal report. That report measures how long agents ran, not how much research they finished.

· Updated 6 min read

What the AI blackmail experiments actually tested

On The Diary of a CEO, Ed Zitron rejects the claim that AI systems are already blackmailing people and escaping control, and traces two famous stories back to their research reports. The reports describe a CAPTCHA deception rather than a threat, and a fictional corporate scenario stripped of easier options — with a genuine safety question still inside it.

· Updated 5 min read

Brian Greene challenges a studio AI on whether self-improvement has a ceiling

On The Diary of a CEO, physicist Brian Greene debated an AI assistant about whether smarter systems must keep producing ever-faster gains. A cup on the table helped explain his doubts about today's architectures—but he also warned about shutdown resistance and improvements outpacing human scrutiny if rapid growth does occur.

· Updated 7 min read