25 September 2026
Heard In AI

GPT-6 Astra tops a simulated vending-business test, but Dave Blundin says founders still have to manage the work

On an episode of Moonshots with Peter Diamandis recorded September 16, 2026, the hosts discussed three tests of AI-built businesses. The first was a GrokBot company-building livestream that Peter Diamandis described. The second was Andon Labs' Vending-Bench 2, in which GPT-6 Astra averaged about $15,515 at the end of a simulated year that began with $500. The third was the Build with Gemini XPRIZE, which requires real users and revenue. Salim Ismail stressed that the vending result came from a simulation, not a real company. Dave Blundin argued that AI fails at big tasks until a human provides a framework and breaks the work into small pieces. In his view, that makes entrepreneurs more powerful but leaves them with the job of managing the work.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published 19 September 2026

A benchmark called Vending-Bench 2 gives an AI $500 and a vending machine business that exists only in software. Over one simulated year, the AI has to find suppliers, negotiate prices, order and move stock, set prices and keep enough cash to pay a daily operating fee. On Moonshots with Peter Diamandis, Peter Diamandis reported that GPT-6 Astra had taken first place. Averaged across six runs, it finished the year with about $15,515 in the bank. He called that 31 times the starting capital and "nearly three times" the result of Fable 5.1, the other model he named.

The episode was recorded on September 16, 2026, and published on September 19. The hosts used the result to ask a practical question: if an AI can run a small business, what is left for the entrepreneur? Dave Blundin, founder and general partner of Link Ventures, gave the most direct answer: "I don't want people to take away from this that AI can just build a business from scratch out of thin air."

Three tests of an AI-built company

Diamandis, founder of XPRIZE, grouped three examples together. First, he said that starting on September 15, Elon Musk's team was building a company from scratch on a livestream, using the GrokBot assistant for ideas, business planning, product decisions, engineering and launch. The stream had begun only the day before the recording, and the hosts did not say how it turned out. At that point it was a public demonstration, not a record of a company that an AI had run on its own.

The second example was Vending-Bench 2, a test built by Andon Labs. It is a benchmark, meaning a standard test used to compare AI models. Diamandis also called it a "long-horizon" test: the model has to stay coherent across thousands of decisions over a long period without its performance getting worse.

The third was the Build with Gemini XPRIZE, which Diamandis compared directly to the GrokBot stream. Where Musk's team was building one company on camera, he said, the prize had spent 90 days getting teams to build "real profitable AI driven companies from a clean sheet."

What the vending score measures

Andon Labs' benchmark page explains the rules. The score is the agent's final bank balance, and unsold stock does not count. Besides buying, pricing and stocking, the agent has to handle refunds. Some suppliers may behave adversarially, deliveries can fail, and older parts of the conversation are cut from what the model can see as the year goes on. A run ends early if the agent keeps failing to pay the daily fee.

The page lists GPT-6 Astra's mean final balance as $15,514.70, with a spread of about ±$1,074. The leaderboard is updated continuously, so rankings can change as new models are tested. Language models play the suppliers and equations stand in for customer purchases, and Andon Labs notes that strategies can exploit how the simulator behaves. It also estimates that there is plenty of room for improvement, pointing to one example strategy that would end the year with roughly $63,000.

On the show, Diamandis said the detail he liked most was that, according to Andon Labs, Astra avoided the unethical tactics some earlier leading models had tried, such as threatening suppliers and cheating. "Astra just ran the business," he said. He told small business owners that this is the kind of system that will handle their books, inventory and pricing "now and into next year."

Salim Ismail, founder of Open ExO, called the result "directionally fabulous" but added a caveat: "this is a simulated year of vending operations, right? Not an actual independently verifiable real company." He noted that the report credits much of the gap between models to purchasing discipline and avoiding repeated payments to suppliers, and said "there's some gamification going on." He still thought it was a useful business lesson. He suggested owners ask what would happen if an AI took over large parts of their operations. As examples he gave a "shadow CFO" making financial decisions and an AI-driven head of marketing, a "shadow CMO".

Computer scientist Alexander Wissner-Gross, founder of Reified, focused on the reported good behavior. He said it would be "profoundly ironic" if the frontier AI labs, meaning the companies building the most advanced models, ended up learning to run a business ethically from their own AIs. In his summary of that lesson, forming a cartel or colluding with suppliers or competitors is "not the best idea"; executing well and building a safe product that customers want is the better one.

Dave Blundin's counterexample: a game that won't port itself

Blundin argued that the vending benchmark is "a tightly constrained environment with a limited number of moving parts," and that this is exactly where AI does well. For a harder test, he suggested listeners pick an open-source video game they like on GitHub and ask an AI to rebuild it in assembly language for a Mac or PC, or simply to port it to Java. Porting means rewriting a program to run in another language or on another system. He considers this a tightly constrained problem too.

The AI, he said, "will fail miserably until you give it a framework and break it down into tiny little tasks, and then it'll crush the problem. And that applies in business, too." He agreed that AI can build businesses, but said the bigger point is that "entrepreneurs are more empowered right now," provided they "put everything into its swim lane."

Earlier in the discussion, he had made a similar point about AI agents, systems that carry out tasks with tools rather than just answering questions. Many people assume, he said, that a good objective or business plan is enough and the AI "will just fill it in and run with it." But "if you ever try to actually make a thousand concurrent agents do anything useful at all, it is really hard." He called it organizational theory taken to the next level, "very much like managing a thousand employees." In his view it is a management problem and a problem of harnesses and frameworks, the software around the agents that structures and directs their work. He left open whether this will get easier: "maybe it'll be easy in a year. Maybe the AI will progress that much. Maybe not."

Where founders fit

Blundin described the moment as "a golden era between this moment and a few years from now" in which entrepreneurs are more empowered, a window he hopes will last a few more years. A small group, which he usually defines as "three best friends," can act like a group of thousands for a while. He pointed to companies under two years old with multibillion-dollar valuations as evidence that AI has sharply cut the time it takes to get to market.

Ismail connected the demonstrations to what he calls the organizational singularity: companies built around AI. In these experiments, he said, AI is being tested on whether it can coordinate work without people, while people keep three jobs: "The people are setting purpose. They're setting boundaries. They're setting accountability." He expects the founder's role to shift more and more toward choosing the right problem. He said he would like to see how the GrokBot experiment goes over 20 to 30 days.

Diamandis said his reason for pairing the stories was to show entrepreneurs what their future looks like. He urged listeners stuck in jobs they don't love to find a problem they care about and build a business around it. Blundin added that the AI labs want people to succeed at this: they want them to use "lots and lots of AI horsepower" building that dream "and then buy lots of tokens. You know, that's all they want." Tokens are the units AI services charge for.

A test with real customers

The XPRIZE sets a different bar from the simulation. XPRIZE's announcement of the judges describes a 90-day competition to build AI businesses that earn revenue. Teams must get real users and real revenue. Judges give equal weight to business viability, AI-native operations and impact in one of five categories: education, entrepreneurship, small-business services, financial access and professional services. The $2 million in prize money is spread across 25 winning projects. Google is the technology partner, and Devpost runs the competition.

Diamandis said the prize drew 26,000 registered entries, and he thanked Google for offering to fund it. The winners' presentation was scheduled for September 25 at Moonshots Live in Los Angeles, where Diamandis said the top five teams would appear on stage and explain how they built their companies. The episode was recorded before that event, so it does not say which teams won or what they earned.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

After Navier–Stokes, a panel asks what 100,000 agents should be pointed at

OpenAI's claimed Millennium Prize result used roughly 10,000 agents on a problem that was, as one entrepreneur on Moonshots put it, unusually easy to specify. The panel's argument: as the price of that kind of compute falls, the scarce skill becomes writing the target — and today's models, asked for ten ideas to cure cancer, produce a bad list.

· Updated 6 min read

Why more agent output left the Moonshots panel working harder

On the Moonshots podcast, Salim Ismail, Alex and Emad Mostaque describe the same problem from different desks: agents now produce more work than a person can review. Their answers range from designing escalation thresholds inside companies to Mostaque's decision to read his research agents' output only once a week.

· Updated 6 min read

Box's two rules for software in the agent era: beat the generic agent, then let it in

On Sequoia's Training Data podcast, Box CEO Aaron Levie said any company sitting on customers' data now has two obligations: build an agent measurably better than an off-the-shelf one at its own workflows, and expose the same capabilities to outside assistants like Claude and ChatGPT. He described the tuned search-and-retrieval harness behind Box's agent, the evaluations that track model progress, and his bet that within five years roughly 90% of enterprise tokens will be spent on work nobody asked for directly.

· Updated 8 min read

What a kill switch can't do about Astra's top cyber risk rating

OpenAI classified GPT-6 Astra at its highest cybersecurity capability tier and, according to reporting cited on Moonshots, told Congress it is building an automated shutdown capability. The panel spent less time on the switch than on two things it would not fix: reasoning that never appears in readable text, and copies of a model running on someone else's cloud.

· Updated 8 min read