October 4, 2026
Heard in AI

Wissner-Gross questions cost case for Sonnet 5.5 and Gemini 4 Argon

Alex Wissner-Gross said Sonnet 5.5's 70% Terminal-Bench score is no better value per attempt than Opus 5.5; Anthropic's charts put the savings at lower effort settings. He said Gemini 4 Argon stands out mainly on avoiding hallucinations.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published October 3, 2026 (recorded October 2, 2026)

Anthropic's new model Claude Sonnet 5.5 posted a large jump on a demanding agent benchmark. Alex Wissner-Gross still said he has no plans to use it. "It's not at all obvious to me why anyone should be using Sonnet 5.5 over Opus 5.5, unless you have some token or latency or other consideration," said Wissner-Gross, a computer scientist and the founder of Reified.

He said this on an episode of Moonshots with Peter Diamandis that was recorded on October 2, 2026, and published the next day. The panel looked at the week's two big model releases: Sonnet 5.5 and Google's Gemini 4 Argon. Wissner-Gross judged both mainly by how much capability each one delivers for the money. The talk then turned to why that question still matters. Richard Socher, who runs the search company You.com and the startup Recursive, said the computing power these models run on is getting more expensive, not cheaper.

A big jump on the benchmark

Host Peter Diamandis, founder of XPRIZE and Singularity University, laid out the headline numbers. Terminal-Bench 4.0 measures how well an AI agent can do real work at the command line, the text-only interface programmers use to control a computer. On that test, Diamandis said, Sonnet 5.5 "jumped from 10% to 70%" in a single generation. It also beat Anthropic's own top model, Opus 5.5, which scored 66.4%, at half the price.

Anthropic's launch announcement matches those figures. It reports 70.6% for Sonnet 5.5 against 10.3% for the previous Sonnet 5, and 66.4% for Opus 5.5. The "half the price" claim refers to per-token list prices. AI companies charge by the token, a small chunk of text. Sonnet costs $2 per million tokens of input and $10 per million of output, while Opus costs $4 and $20.

The benchmark itself was also new. Its maintainers say version 4.0 gives every task an eight-hour time limit, recalibrates the computing resources each task gets, removes eight tasks and fixes nineteen others. Because the tasks and test environments changed, they say results from earlier versions cannot be treated as interchangeable with 4.0 scores. Anthropic reports both Sonnet numbers on version 4.0.

"Another really weird release"

Wissner-Gross said his doubts came from the launch chart, which plots each model's score against what it costs to get that score. A model's "cost-performance frontier" is the best score it reaches at each level of spending.

He explained the pattern he expected. Labs usually release a large model and then a smaller one that has been distilled from it, meaning trained to imitate the bigger model more cheaply. On a chart like this, the smaller model normally sits "up and to the left" of its parent: an equal or better score for less money, or what he called "greater intelligence per dollar." Sonnet 5.5 didn't do that. In his reading, its line, drawn in blue, was simply "a visual extrapolation" of the red Opus 5.5 line. On a cost-per-attempt basis, he said, Sonnet 5.5 scores lower than Opus 5.5.

He granted the obvious point: Sonnet 5.5 is "a big jump over the past Sonnet." He also allowed that someone could argue it is better when measured per token. "But on a cost performance basis," he said, "it's actually worse, it appears, than Opus 5.5."

Anthropic's own account makes the answer depend on a setting the user chooses. The company's charts compare cost per task at different effort levels, meaning how much extra work the model does on a problem before it answers. Anthropic says Sonnet 5.5 at low or medium effort can beat Sonnet 5's best score at about one-tenth of its task cost, and it describes lower-effort Sonnet as a complement to Opus. At higher settings, the company says, Sonnet can reach comparable performance at a similar cost. That part fits Wissner-Gross's reading. Anthropic also pitches Sonnet for bounded everyday work such as bug fixes and document tasks, and recommends Opus for more complex work that needs sustained judgment.

Diamandis also noted that Sonnet 5.5 is the first Sonnet model Anthropic launched with cyber safeguards. According to the announcement, these include sending certain high-risk requests to the older Sonnet 5.

Dave Blundin's theory: keeping customers "top to bottom"

Dave Blundin, founder and general partner of Link Ventures, offered a business explanation for the pricing. He said many enterprises, the large companies buying AI, are moving toward orchestration. In this setup, an expensive model acts as a manager that splits up a job and hands the pieces to cheaper models. The sales pitch he described is to use Opus 5.5 as the orchestrator and models such as Kimi K3 or Qwen as sub-models "at, you know, half or a third or fifth the price." That can roughly halve the cost of each finished outcome, he said, but "it relies on Anthropic up here and cheaper models down here."

By cutting the cost in half, Blundin suggested, Anthropic may be trying to fill the bottom of that stack itself and tell customers to "go with Anthropic top to bottom," with no need to worry about "Chinese code injection." His theory was that Anthropic is competing with China, "which is about three months behind," and trying to close that gap "before a lot of enterprises go to open source models." He said Alex Karp is pushing the opposite message: companies that want to control their own destiny should not get addicted to one vendor and should use models they control.

Salim Ismail, founder of Open ExO, said the bigger shift is how quickly any model loses its edge. Models have "an increasingly shorter and shorter half-life," he said, and since everybody has access to the newest one, the advantage comes from how an organization can "metabolize that into some decent capability."

Gemini 4 Argon: third place, and few hallucinations

Google announced Gemini 4 Argon on September 30. It is built for long-horizon work, meaning tasks that run over many steps, in coding, legal and financial work, and defensive cybersecurity. Its output limit rises from 64,000 tokens to one million. Diamandis said Google had promised a frontier model, Gemini 3.5 Pro, in June, but it was delayed internally and never shipped, and he called Argon the company's comeback. Google says the first users will be trusted cyber defenders, with wider access planned after testing.

Wissner-Gross started with a disclaimer that he has friends on every lab's model team, "but I view one of my jobs here is to call objectively balls and strikes." Blundin guessed this one would be a ball. Wissner-Gross later joked that he "used to, until right this moment, have friends on the Gemini team."

His verdict: Argon does not put Gemini at the cost-performance frontier or at the capability frontier. It does put Google back in the top three labs, behind Anthropic and OpenAI, but not in the top two. He called the benchmarks Google chose to highlight "mildly cherry-picked." He pointed instead to the Artificial Analysis index, which combines several capabilities and ranks Google third. If you draw the line connecting the best-value models on cost against performance, he said, Argon "doesn't even make the optimal frontier."

The one area where he said Argon is "beating the pants off of everyone else" is avoiding hallucinations, the confident but made-up answers language models sometimes give. He offered what he called a "Kremlinological analysis," reading the company from the outside. In his view, the Gemini team has "two masters": outside developers, and the answer box at the top of Google search results, which is powered by a Gemini model. Google is "presumably burned by past experiences," he said, and doesn't want that box to give wildly wrong answers. He linked this to internal competition for scarce chips. As far as he can tell, including from chatting with people at Google this past week, he said, it is still "an internal knife fight" over computing power between Google Cloud, which wants to sell it to outside customers, search and ads, and Google DeepMind, which needs it to train and run models.

Blundin played with the name: argon is an inert gas, so a model whose biggest strength is not hallucinating is "a pretty inert gas model." ("Nominative determinism," Wissner-Gross replied.) More seriously, Blundin said a company that has already lost a major antitrust case and is going through the settlement process faces big liability risks when it puts a model in front of every search user. "The fear way outweighs the opportunity in the mind of the big corporate giant," he said.

Socher, whose company You.com competes with Google in search, said whether hallucination is good or bad "depends totally on the context." In drug and protein discovery, you want the AI to come up with new ideas, such as "novel combinations of amino acids to create new proteins." In a search engine, you usually want none. He said Google and others took a long time to catch up with You.com on accuracy and citations, even though You.com has far fewer resources. In his view, Google has realized its business doesn't necessarily need superintelligence: people come to it for quick questions like a good restaurant, not to "solve the Riemann hypothesis for me."

Why price per task still matters

Diamandis asked Socher why computing power, which Socher had called the biggest constraint, still holds things back when intelligence keeps getting cheaper per task. "Just, the physics, I guess, of it," Socher said. If you try to buy a thousand of Nvidia's GB200 chips, he said, "the price has actually gone up, in several cases." Even H100 chips, which are seven years old and which financial models usually assume are worth nothing after five, have gone up in price "in a crazy way" over the last few months, he said.

Blundin had a story of his own. He said he had an eight-chip B300 system on order for $3 million, due in December, and then got a call: it had been sold to someone else for $5 million, "$2 million more." He later said it was actually an NVL72 rack.

Socher called the situation "a bit of a compute crunch" and said it is something "capitalism will solve." Demand is so high that many people are building land, power, building shells and data centers. His hunch is that in about two years more capacity will reach the market, and computing will then be priced "a little bit like electricity," going up and down. He doesn't expect a crash, but for now buyers of the largest GPU clusters are trying to lock in prices because they expect them to keep rising for the next few months.

Blundin then asked whether Socher had just spent $450 million. Socher replied that it "may have been one of the smallest compute deals we've done."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push | EP #299

Episode published (recorded )This article draws on 0:18–0:45 and 1:35:30–1:50:08 (approximate times)

Article history

Updates to this article

Tags