27 September 2026
Heard in AI

GPT-6 Astra's cone-course drive splits Moonshots over specialist AI

On DrivingBench, a new benchmark, OpenAI's general-purpose GPT-6 Astra drove a real Toyota Corolla through a whole cone course on its second attempt. No other model tested finished the course. On Moonshots, Alexander Wissner-Gross argued that this makes specialised robotics models unnecessary. Dave Blundin said a second-try pass would not satisfy drivers, and he asked how much computing power the drive used. The "100%" score measures how far along the course the car got, not how reliably it drives. Each model had only one session, and the models ran through different software tools.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published 27 September 2026 (recorded 25 September 2026)

A general-purpose AI model has driven a real car through a course marked out with cones. It is the same kind of model people use to write code or answer questions. On a Moonshots with Peter Diamandis episode recorded on 25 September 2026, the panel disagreed about what that proves. Alexander Wissner-Gross, a computer scientist and founder of Reified, said it shows robotics no longer needs its own specialised models. Dave Blundin, founder and general partner of Link Ventures, said the result leaves open the questions that matter for putting such systems to work: reliability and cost.

What the car actually did

Peter Diamandis, founder of XPRIZE and Singularity University, introduced the benchmark, called DrivingBench. A group of developers fitted a Toyota Corolla with an interface that lets software steer, accelerate and brake. They then gave a general-purpose model a single instruction to drive from one point to another through what Diamandis called a 130-metre cone course. "GPT-6 Astra finished the entire course 100% on its second try," he said.

The DrivingBench report explains how the test worked. The car is a 2022 Corolla controlled through comma hardware and the openpilot driving software. The model does not see the road as a person would. It receives camera frames and telemetry, meaning live readings from the car's sensors, and it replies with motion commands kept within set limits. Each model got one session of up to three attempts. It kept its conversation history between attempts and could reflect on what had gone wrong before trying again. A human safety operator stayed ready to brake throughout.

The "100%" does not mean Astra drives safely every time. According to the report, the score measures progress along the centre line of the course, and progress counts only while the car stays within four metres of that line. Driving farther in the wrong direction earns nothing. The leaderboard shows Astra reaching 49% on its first attempt and 100% on its second, which took five minutes 22 seconds. No other model finished. Claude Fable 5.1 reached 9%, 10% and 45% over its three attempts. Grok 4.6 reached 8%, 11% and 10%, and GPT-5.6 Sol reached 6% each time. The authors say the other models often failed because they misread where the cone boundaries were.

The leaderboard's comparison is also less even than it looks. All models used medium reasoning settings, but each ran through a different application harness, the software that connects a model to a task: Codex, Claude Code or Cursor. Each model had just one session, so the result comes from a single run rather than a measured success rate. The authors list further limitations: limited steering range, camera blind spots, overshooting the starting speed and no independent repeated sessions. They propose repeated evaluations, more reasoning settings and harder courses.

"Bitter lesson is bitter indeed"

Diamandis read out a reaction from Boris Power: "Incredible results. This should be a nail in the coffin of specialized models that were trained from scratch, with a lot of specialized effort, versus just training the most powerful generalized model, and eventually distilling a small specialized model as needed." Distilling means training a smaller, cheaper model to copy what a larger one can do.

Wissner-Gross went further. He opened with the line that the "bitter lesson is bitter indeed." The phrase refers to computer scientist Rich Sutton's argument that general methods which use more computation eventually beat systems built on hand-crafted expert knowledge. Wissner-Gross said that, "either now or imminently," building a robot could become an elementary-school project. "You'll just vibe code a robot," he said. Vibe coding means describing what you want and letting AI write the software. You would then drop a frontier model such as Astra into a robotic body, and he said it "will one-shot embodied cognition". In other words, it would work out how to act in the physical world on its first try.

He expected this to be "pretty upsetting to a number of academic computer scientists." With hindsight, he said, decades of work in computer science and robotics were probably a "total waste". In his words: "All we needed was a generalist model." Diamandis joked that superintelligence being able to drive a car was hardly news: "What a surprise."

Diamandis then asked Emad Mostaque, founder of Intelligent Internet, for his view, and Mostaque broadly agreed. He argued that a general model's strength will carry over to biology, to physics and "to just about anything." "I don't think there's a single specialized model that you can say will outperform in a couple of years' time," he said.

Capability is not the same as a usable driver

Blundin accepted part of the argument. "I think both things can be true," he said. A general model can drive a car, act as a robot and even build one. But he separated that from being good enough to use. "If you said, hey, it completed the course on its second try, Elon's going to be like, yeah, your Tesla can't do that on a second try," he said. "That's not going to work for most drivers."

His larger objection was about efficiency. "The thing that's missing in the storyline there is how much compute did you use to do the task?" The DrivingBench leaderboard does list the tokens used and the cost at list prices for each attempt, including the reflection that follows it. Astra's successful second attempt used 6.6 million tokens and cost $7.74. Its first attempt used 1.2 million tokens and cost $2.01. Those figures show what it cost to use the model. They do not measure the chips, memory or electricity behind the drive, and the benchmark does not report those. For Blundin, the amount of compute used decides whether a general model makes sense for a job. "Compute is going to be forever starved from here forward," he said. "If you used more compute to do the same task, that's a crime." He did not dispute that a general model can do the task. His question was what it costs to run one.

From models down to memory and energy

Blundin tied the scarcity of computing power to hardware prices. He said chip prices have risen and that "HBM RAM has gone up 5x in price." He blamed demand: "the AI is so incredibly valuable, people want to compute." HBM, or high-bandwidth memory, is the specialised memory that sits beside AI processors. Micron explains that HBM stacks memory chips vertically and links them to the processor through a wide connection over short distances. That lets it move more data in parallel, using less energy for each bit it moves. For an AI chip, extra calculating power helps only if memory can supply data fast enough. The fivefold price rise is Blundin's figure.

Salim Ismail, founder of Open ExO, called this "a really important point". In his view the bottleneck is shifting "from the models down the stack to compute, and then eventually energy." Diamandis said the industry had already reached the energy stage. "We're short, like, 60 gigawatts of energy in the next two years," he said. If general models take over the work of specialised ones, the panel's view was that the limit becomes the chips, memory and electricity each task consumes. DrivingBench's list-price figures show what each attempt cost, but not how much of that hardware and energy it used.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03

Connected ideas and articles

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Why Jensen and Zuck think the doomers are wrong (plus AI get’s a rebrand) | #294 MOONSHOTS Live

Episode published (recorded )This article draws on 39:53–41:46 and 42:34–43:54 (approximate times)

Article history

Updates to this article

Tags

DeepSeek's leaner cache changes data center buys, Moonshots panel says

On Moonshots #288, a 4 a.m. chart about DeepSeek's new V4.1-Flash model sent the panel from cache statistics to the shopping list for an AI data center. DeepSeek says the model's lookup memory needs a quarter of the expensive high-bandwidth memory and an eighth of the SSD cache storage of its previous generation. The panel's argument was about what that does to a buildout in which, by one panelist's estimate, 40% of American capital spending goes to that one component.

7 min read

Moonshots panel on why China's AI tokens go to video, US's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

6 min read

After Navier–Stokes, a panel asks where to point 100,000 agents

OpenAI's claimed Millennium Prize result used roughly 10,000 agents on a problem that was, as one entrepreneur on Moonshots put it, unusually easy to specify. The panel's argument: as the price of that kind of compute falls, the scarce skill becomes writing the target — and today's models, asked for ten ideas to cure cancer, produce a bad list.

6 min read

What OpenAI's 10,000 agents proved about Navier–Stokes blowup

OpenAI said on 8 September that an internal model, running roughly 10,000 agents for 88 hours, produced a forced blowup construction for the Navier–Stokes equations and a machine-checked proof of it. On Moonshots with Peter Diamandis, the panel worked through what the result is — a statement about idealized fluids, not a device — what it cost, and why the credit for it was contested within hours.

7 min read