October 1, 2026
Heard in AI

How to read GPU and token prices, per Silicon Data's Steve Hou

Steve Hou of Silicon Data said big clouds charge at least two to three times neocloud GPU rates for bundled services. He guessed a recent rise in its token index reflects which models people use, not price hikes.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 1, 2026

Renting a GPU from one of the big cloud providers usually costs at least two to three times what a smaller specialist charges for what looks like the same chip, according to Steve Hou. In his view, the gap does not show the market getting the price wrong. Hyperscalers sell the chip as part of a different product.

Hou is head of research at Silicon Data, which builds price indexes for rented GPU capacity and for AI model tokens. He explained how those indexes are built, and how to read them, on The Cognitive Revolution in an episode published October 1, 2026. Host Nathan Labenz and co-host Prakash Narayanan asked the questions.

Some background. GPUs, or graphics processing units, are the chips that train and run most AI models. Companies often rent them by the hour instead of buying them. Sellers fall into two groups: the hyperscalers, meaning the giant general-purpose cloud providers, and "neoclouds," smaller providers that mainly rent out GPU capacity. A token is the small chunk of text that language models read and write, and AI companies price their models per million tokens.

Quoted prices versus paid prices

Hou said Silicon Data's GPU index uses both quoted prices, meaning what sellers advertise, and transaction prices, meaning what renters actually paid. It labels which is which. It then runs what he called a large normalization process, using machine learning to make contracts "apples to apples" comparable. The aim is that an H100, a widely used Nvidia data-center chip, rented in one part of the world can be compared with one rented somewhere else.

The company's methodology write-up describes the scale: about 150,000 verified daily pricing records from 50 to 100 platforms in 40 to 50 countries and regions, with coverage starting September 1, 2024. The contracts range from on-demand capacity and interruptible "spot" rentals to reservations lasting from one to 60 months. Silicon Data says simply averaging a short spot rental with a multi-year reservation could make a change in the mix of contracts look like a change in market prices. Normalizing prices to a standard configuration is meant to prevent that.

Hou expected one objection to using advertised prices. Nobody building an apartment rent index, he said, would trust the number posted on the front of a building. GPUs are different, in his account, because "you cannot click an API and just get hold of the apartment." Much GPU rental can be done through an API, a programmatic interface that lets software place an order directly. That makes a quote close to a real offer, "the same way you can buy something on Amazon."

Transaction prices have a weakness of their own. If a seller rented out three nodes at one price, Hou said, he would not assume a fourth would be available at that price. If everyone else is quoting $4, a seller still renting at the old, lower price is either about to raise it or there is "something wrong with that price." Combining both kinds of data is his way of capturing "the market as it is, as faithfully as possible."

The burger in the steakhouse

Labenz said that, just browsing Silicon Data's website, he had noticed hyperscaler prices running "seemingly 2 to 4x" those of neoclouds. He asked why the gap exists and why nobody arbitrages it away.

Hou answered with a burger. Imagine a single-patty burger index, he said. A burger from a New York street stand and a burger in a high-end steakhouse will cost very different amounts. The index tries to find the price of the same basic unit with the same features, so it would not compare a three-patty burger with a single-patty one. Some adjustments, though, he said he "cannot reasonably make" because the price includes a premium for bundling or differentiation.

Hyperscalers, Hou said, charge "regularly, consistently, at least two to three times, sometimes more" than a typical neocloud. He put that down to their long history: other products on the same platform, software analytics, safety and compliance, and long-standing relationships with enterprise customers who are reluctant to move. "It is being sold as a very much of a differentiated product," he said.

He took the burger analogy further. If a seller lists a burger, fries and a milkshake as separate items, he can subtract the extras and isolate the burger's price. But if the fries and milkshake were blended and "inject[ed] into the patty" and sold as a new product at two or three times the price, they cannot be stripped out. At that point, Hou said, he puts up his hands and files the product in a different category. That, he said, is how Silicon Data has handled hyperscalers so far. Separating them gets the index closer to a pure unit of compute, one where offerings are direct substitutes for each other.

Silicon Data's own regional analysis from July 31, 2026 offers a somewhat smaller figure. Comparing median H100 rents in seven regions since early 2025, it found the hyperscaler multiple had narrowed from roughly 3 to 4.5 times to an average of about 2 to 2.5 times, ranging from about 1.6 times in the Nordics to 3 times in US West and the Middle East. The current quarter in that data was incomplete. The analysis lists similar reasons for the premium: bundled orchestration, enterprise support, compliance, security and integration with existing storage and networking.

The GPU lottery stays out of the price

Narayanan noted that the same chip can perform differently in different places, and asked whether every price in the index is checked with Silicon Data's benchmarking software. Hou said no. The company does run a physical benchmarking service, SiliconMark, that tests individual GPUs for health, throughput and other measures. But that physical performance, he said, "does not enter into our pricing," because for now the company sees no strong relationship between how a particular GPU performs and its price.

He acknowledged that chips do vary, which people in the industry call the "GPU lottery." A paper by Silicon Data researchers shows what that means. Across 3,534 GPUs from 11 anonymized providers, most individual devices gave consistent results when retested. Measured across devices and providers, though, the spread was much larger: one H100 variant's computing throughput varied by as much as about 34%.

Hou expects that variation to even out over a cluster of many chips, where "a law of large numbers kicks in." The index instead uses six features, including location, CPU memory, contract term and provider. He said physical performance could become a pricing factor later. His thesis is that inference, meaning running trained models instead of training them, could make compute more interchangeable, and that could eventually lead to markets with physical delivery of compute.

Why a token index can rise while prices fall

Labenz's other question was about tokens. Silicon Data's proprietary-model token index showed a rise of 11.5% over seven days. Labenz pointed out that, adjusted for quality, everybody would say model prices are falling, and that posted API prices had not changed except when new models came out. So what was the index measuring?

Hou said the indexes have been "repeatedly" misunderstood, partly because of what he called "the unfortunate naming." Calling it an expenditure index led people to think it tracks either price or total volume. It is neither. It is an expenditure-weighted price per million tokens. The index documentation describes it as a usage-weighted price that combines providers' listed prices with activity observed on routing and inference platforms. It is not total spending and not a quality-adjusted measure of intelligence, and private discounts can make what a given customer actually pays different.

Hou put the rise in context. Since around late June, he said, the proprietary index had dropped sharply, from about $4 to about $1.60, "more than a 50% drop." On those figures the fall is roughly 60%. During that time, leading labs released cheaper versions of powerful models, and other companies with proprietary models, which he named as Meta and Grok, competed aggressively on price.

To explain the recent rise, Hou used cars. If an automaker like Mercedes sells cheap and expensive versions, the average price of cars sold depends on which versions buyers pick. Tokens work the same way: model prices change, and so does what people use. If more users switch to an expensive, powerful model from a company like Anthropic, the weighted index can rise even if no price list changes. The documentation agrees that a shift toward pricier models, new model releases, or a change in the balance of input and output tokens can move the index.

Hou said he had not looked into the details of the latest rise. His guess was that it "has more to do with usage mix than anything else."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04

Connected ideas and articles

From the conversation

Podcast episodes

The Cognitive Revolution

AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases

Episode published This article draws on 0:51–1:34 and 31:40–42:00 (approximate times)

Article history

Updates to this article

Tags