October 9, 2026
Heard in AI

Positron's CTO makes the case for a memory-first AI chip

Positron CTO Thomas Sohmers said AI agents are multiplying demand for memory. He claimed Positron's first-generation product sustained 93% of theoretical memory bandwidth, against an average of 30–40% for NVIDIA GPUs when decoding with transformer models. Its next chip, Asimov, is planned for 2027.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 8, 2026

A year before his interview, Positron had a price quote for memory. When Thomas Sohmers, the company's co-founder and chief technology officer, looked at it again, he found that the price had risen four and a half times. "I really hope that this isn't the case," he said, "but I wouldn't be surprised if this is going to end up being another doubling over the next year."

Sohmers spoke on a weekly highlights edition of The Cognitive Revolution, published October 8, 2026. Host Nathan Labenz introduced Positron as a company that builds chips for running AI models and said it had recently raised $875 million. Asked whether the price rise had changed the economics, Sohmers said people have absorbed the higher costs because the same memory now produces much more. A model running on a given number of gigabytes, he said, is "way more than 5x" as capable as it was a year earlier.

Why memory is the bottleneck

Running a trained model to answer requests is called inference. Inference needs memory for two things. The first is the model's weights, the billions of learned numbers that make up the model. The second is the context for each user or session: the conversation, documents and intermediate state the model has to keep available while it works.

Sohmers said he founded Positron believing memory would become the limiting factor, even though people still argue about power and other parts of the infrastructure. He sees no near-term end to scaling, the pattern in which larger models deliver more value, and larger models take up more memory for their weights. He also said the main obstacle to using AI in most applications today is context length: holding more context for each user, then serving many more users.

Agents add to that demand. These are AI systems that work through tasks on their own, often in the background. Sohmers gave his own case as an example. Three or four months earlier, he usually had two to four agents running constantly. Now he runs 15 to 20. Each agent keeps its own separate context in memory. Multiply that by everyone using AI, he said, and the total "adds up very quickly."

That reasoning led Positron to build on commodity memory, specifically LPDDR, the kind found in phones and laptops. "We have to use the commodity number," Sohmers said, "because that's the only thing that's going to be able to scale and be cost effective."

Using the bandwidth you pay for

Sohmers said the design response had two parts. The first concerns memory bandwidth, the rate at which data can be moved from memory to the processor. This matters most in decoding, when a model generates text one token at a time. A token is a small chunk of text, and producing each one requires reading through the model's weights.

Sohmers said NVIDIA GPUs average 30 to 40% memory bandwidth utilization during that decode step. On a B300, which he said is advertised at 8 terabytes per second, he put the realized figure in the ballpark of 2 terabytes per second. He blamed GPU architecture: the hardware was designed around data-reuse patterns that suit training and other workloads but "really don't exist in transformers," the architecture behind today's language models. He said Positron's first-generation product reached and sustained 93% of its theoretical bandwidth, which he called roughly a 3x improvement. These utilization numbers are his own claims. For its next chip, Positron's Asimov product page lists 2.76 terabytes per second of realized memory bandwidth.

The second part is the number of memory channels, the separate paths connecting memory to a chip. Sohmers said other products top out at about 12 to 16 LPDDR channels. Positron goes up to 72 channels of LPDDR5X through what he described as a decoupled memory chip solution developed with Credo Semiconductor. The likely component is Weaver, a "memory fanout gearbox" that Credo announced on November 3, 2025. It separates the accelerator's connections from the memory interfaces so that more memory can sit around one chip. In that announcement, Positron CEO Mitesh Agrawal named Weaver as part of the company's next-generation inference servers. Credo said it planned to make Weaver available in the second half of 2026. That is a schedule, not confirmation of delivery.

One chip instead of eight

Positron's next chip is called Asimov. Sohmers said it has "up to 2.3 terabytes of memory capacity per chip," compared with 288 gigabytes on NVIDIA's B300. NVIDIA's DGX B300 documentation confirms 288 GB per GPU. It describes an eight-GPU system with about 2.3 TB of GPU memory in total. So Sohmers's figure for a single Asimov chip matches an entire eight-GPU NVIDIA server, which is how he described it: what would take eight GPUs for memory can fit on one device.

The Asimov page gives the per-chip range as 288 GB to 2.3 TB and describes the chip as planned for 2027. The 2.3 TB figure is the top configuration of a product that has not yet shipped. The B300 comparison, by contrast, is with hardware NVIDIA ships today.

Sohmers said the gain is more than saving silicon. When a model is split, or sharded, across several devices, the devices must keep exchanging partial results. He named the all-gather and all-reduce operations, which collect and combine data across chips, and said they are needed for every layer and every sharded matrix multiplication. A model that fits on one chip avoids that coordination.

Looping layers saves space, not traffic

The conversation then turned to looping, where a model runs some of its layers more than once for each token. Sohmers traced the idea from a loop-transformer research paper to what he described as a rumor that looping is one of the big advances in GPT-6 Astra. He also pointed to work on the fringes of the open-source community two or three years earlier. People took Meta's Llama 70B model, duplicated some of its layers in sequence and made it something like a 100-billion-parameter model. According to Sohmers, it got better results "even though it was repeating the exact same" computations.

Labenz replied: "Imagine what you can do when you train it to work that way." Sohmers agreed and sketched where training could lead. A model that is already confident about the next token, with several layers agreeing, could exit early. A model that is very uncertain partway through could route back and run the layers most likely to help, much as the router in a mixture-of-experts model picks which specialist sub-networks handle each token. He said he "wouldn't be surprised if these things are already being done at scale in the major model lines," but that is his speculation.

Labenz suggested that the hardware appeal must be bandwidth: if the same weights stay on the chip, less data has to be shuffled in and out. "Not really," Sohmers said. On-chip SRAM, the fast memory built into the processor, is small. By the time a layer finishes, the chip has moved many times more data than that SRAM holds. A second pass therefore fetches the weights from main memory again. "Looping really is more memory capacity savings than memory bandwidth savings," he said.

He gave a worked example. Take two models trained from the same base. Model A has 10 layers and is trained to loop through them twice. Model B has 20 distinct layers. Per token, both move exactly the same number of bytes and do the same amount of arithmetic. The difference is that model B's second group of 10 layers has its own weights, so it is twice the size. If the looped model got, in his hypothetical, 95% of the quality at half the size, "you're probably going to deploy that one." The saving appears in how much memory is needed to hold the model, not in how fast data must move. As Sohmers put it, the 20-layer version means "you require twice the memory capacity."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03

Connected ideas and articles

From the conversation

Podcast episodes

The Cognitive Revolution

AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Episode published This article draws on 0:57–1:16 and 31:13–40:48 (approximate times)

Article history

Updates to this article

Tags