27 September 2026
Heard in AI

Useful AI agents needn't top benchmarks, Blundin says of Meta's Muse

Investor Dave Blundin says Meta's Muse personal agent shows a split in AI. On one side are consumer agents that handle everyday tasks. On the other are frontier models that compete for the top benchmark scores. Blundin chairs EverQuote, whose shares he said fell 15% on back-to-back days. Investors feared Muse would take over insurance and mortgage shopping, which he thinks won't happen. Emad Mostaque argued that models are now smart enough and the priority is avoiding mistakes. Blundin warned that people who start with today's usually correct models may trust answers that are still sometimes wrong.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published 27 September 2026 (recorded 25 September 2026)

When Meta launched Muse, its new personal AI agent, some investors decided that shopping for car insurance or a mortgage online might never be the same. Dave Blundin, founder and general partner of Link Ventures, saw the reaction up close. On the Moonshots podcast, recorded on September 25, 2026, he said stocks including EverQuote, where he is chairman, fell 15% on back-to-back days. In his telling, investors thought: "oh my God, I'm going to always find my auto insurance and my mortgage through MetaMuse."

Blundin said he doesn't think it will "actually play out that way." He still found the reaction revealing. People were starting to ask whether Google search would become irrelevant, he said, because Muse "just looks like my co-pilot for life." That broader shift, he added, "definitely is going to happen."

What Muse is

Meta announced Muse on September 8, 2026. It describes Muse as a personal agent powered by Muse Spark, a model of its own. An agent, in this sense, is software that carries out tasks for you instead of only answering questions. Each Muse agent runs on its own virtual computer with a web browser. It works through connected services to book travel, fill in forms, shop and negotiate, and it keeps working after the user closes the app. Meta says Muse asks for approval before actions such as sending emails or making purchases. Its memory links everyday activities: it can turn a recipe seen on Instagram into a grocery list while remembering the user's dietary preferences. At launch, payments went through Stripe's Link, including single-use card details.

Peter Diamandis, the XPRIZE founder who hosts the show, summed Muse up as an assistant that can book flights, order food, manage a calendar and do the shopping. He said it reached number one on the App Store with 2.8 million downloads, making it in his words the fastest consumer AI product, "surpassing even ChatGPT when it came out." He credited Meta's reach: the company, he said, touches 3.8 billion people, and "if you have that level of connectivity to an end customer, you can push your product quickly."

A fork in the road

Blundin's main point was that the AI business is dividing in two. On one side is consumer AI, which he called "insanely valuable" and expects to "completely replace search and all the other ways you interact with the internet." On the other is the race toward recursively self-improving AI, meaning systems that help build better versions of themselves. That race is the one tracked closely with benchmark tests.

People are used to judging models by those scores, he said, and to calling the latest benchmark leader, such as Anthropic's Claude Opus 5.5, "the coolest model." Muse is doing something different, he argued: it "doesn't have to be the most intelligent model to be the most useful in terms of your day-to-day life."

Why Meta needs it

Alexander Wissner-Gross, a computer scientist and founder of Reified, took a more strategic view: "Meta needs Muse to work." He described Muse as a computer-use agent, an AI that operates a computer the way a person would. He linked it to Manus, a Chinese company of that kind that Meta tried to acquire. According to Wissner-Gross, the Chinese government forced Meta to unwind that deal.

He called Muse Meta's "primary distribution play": a way to use the audience Meta already has to spread its AI. He pointed to Meta using Instagram "to the hilt" to promote Muse and get people to install it. The stakes are high, in his view. Without such a product, he said, Meta "risks oblivion, irrelevance," because Instagram and Meta's other apps won't matter as much once their content is "all synthetic," meaning AI-generated. As he sees it, Meta has to use its existing apps as a beachhead to move people onto its own agent before that happens.

Competence over intelligence

The same idea came up again when the panel discussed the week's new models. Diamandis said Opus 5.5 delivered performance equivalent to Fable 5.1 at "literally half the price," and that he had switched his own assistant to it to save money. Anthropic's announcement is more measured. It places Opus 5.5 near Fable 5.1 on most work and estimates typical costs about 40% lower than those of the previous Opus 5. The roughly half-price comparison with Fable comes from one internal task, rewriting the HAProxy software. Both models' versions passed nearly all tests, and Opus 5.5 finished in 9.5 hours instead of twelve at 51% lower cost.

The conversation then described Opus 5.5 as the first truly competent model, a shift from Opus 5, which was called "completely unhinged." Competence, not a higher score, was presented as what makes products like Muse work.

Blundin gave an example from building software by simply describing what you want, a practice known as vibe coding. When you put a visual interface on a project, he said, the model now shows mock-ups before it builds anything, then builds "exactly what it proposed in the visual." Seeing what the model is doing, and feeling "like you're inside it," is "night and day different" from just days earlier, he said. His advice: "just upgrade all your systems."

The panel moved on to OpenAI's latest GPT-6 models. Blundin said error rates had fallen so far that they undercut an old belief, which he said was "rampant at MIT," that hallucinations, meaning confident but made-up answers, would be with us forever. "We've just scaled out of it," he said.

Diamandis then asked Emad Mostaque, founder of Intelligent Internet, whether he had been using the new models. Mostaque called it "an incredibly competent model" in his own testing. "The models are smart enough now," he said. "We need them not to make mistakes. We need them not to drop the ball." He expects agents such as Muse to do "tremendously well" because they can act as a personal assistant, chief of staff and organizer, "and we all need some competence in our lives."

The catch

Blundin agreed that such assistants are "insanely powerful" as organizers, but he also named the risk that comes with that reliability. People who used GPT-2, 3 and 4 learned to expect misinformation. A child who starts with GPT-5 and 6, he said, is "just going to trust it out of the box because it's right 99.9% of the time." He called that "kind of dangerous because it still does occasionally just give you something factually wrong."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Why Jensen and Zuck think the doomers are wrong (plus AI get’s a rebrand) | #294 MOONSHOTS Live

Episode published (recorded )This article draws on 31:33–36:15 and 38:18–39:53 (approximate times)

Article history

Updates to this article

Tags

Box CEO Aaron Levie's two rules: beat generic agents, then let them in

On Sequoia's Training Data podcast, Box CEO Aaron Levie said any company sitting on customers' data now has two obligations: build an agent measurably better than an off-the-shelf one at its own workflows, and expose the same capabilities to outside assistants like Claude and ChatGPT. He described the tuned search-and-retrieval harness behind Box's agent, the evaluations that track model progress, and his bet that within five years roughly 90% of enterprise tokens will be spent on work nobody asked for directly.

8 min read

What changes when Grok Bot gives each AI agent its own computer

On the Moonshots panel, Peter Diamandis runs a Grok Bot chief of staff called Skippy and Emad Mostaque runs 18 of them across his own machines, installing models and making art. Salim Ismail calls it the move from asking an AI to assigning work; Alex argues the messaging-app interface cannot possibly scale.

6 min read

Why Ed Zitron trusts his editor more than a hallucination score

On The Diary of a CEO, writer Ed Zitron described catching an invented Microsoft share price in his Bloomberg terminal, then argued that his editor Matt Hughes — not a benchmark number — is what makes an answer trustworthy. The host pushed back: buyers pay for the output, not the process, and the honest comparison is AI against fallible people rather than perfection.

7 min read

Dave Blundin says one bad idea can convince all 5,000 of his AI agents

On Moonshots with Peter Diamandis, Dave Blundin, founder and general partner of Link Ventures, described a bad coding idea spreading through a swarm of 5,000 identical AI agents until he intercepts and rewinds them — otherwise, he says, roughly $50,000 of tokens goes into a harebrained plan. The panel set that experience beside a new paper on "mind viruses" that spread between agents through editable memory, and argued about whether "virus" is the right word.

6 min read