October 9, 2026
Heard in AI

Positron co-founder says AI token spend briefly topped salaries

Thomas Sohmers, co-founder of AI chip company Positron, says GPT-6 Astra and Opus 5.5 agents now run its chip tests on emulators in a closed loop. He says the company's token spend briefly exceeded its human salaries and peaked above $100,000 a day, and the agents have not yet had a novel design idea.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 8, 2026

When Positron's engineers handed AI agents "full keys" to the company's internal infrastructure and pointed them at a hardware emulator, the agents did something Thomas Sohmers, a co-founder of Positron, called "mind-blowing." They read the manuals, wrote themselves study notes, built their own testing setups and began checking the company's chip design without a human stepping in at each stage.

Sohmers described the experiment to Nathan Labenz on The Cognitive Revolution, in a weekly highlights episode published October 8, 2026. He also described the bill: by his account, AI usage quickly became Positron's largest expense outside manufacturing, and for a stretch it cost more than the company's human salaries.

Designing a chip versus checking it

Building a chip involves two broad kinds of work. Designing it means deciding what the circuits should do and how they connect. Verifying it means proving, before anything is manufactured, that the design actually behaves as intended. A mistake caught after fabrication can be extremely expensive, so the checking is a large share of the job.

Labenz opened with a comparison from NVIDIA chief executive Jensen Huang, who in an interview with Ezra Klein said NVIDIA spends something like 20% of its effort designing and 80% verifying, validating and ensuring reliability. Sohmers said Positron's ratio looks different because it is starting from scratch and has to build a lot of basic capability itself. For Asimov, its first fully custom chip, he expects the split to end up around 60% design and 40% verification.

Agents at the emulator

For verification, Positron uses Cadence Palladium emulators. Sohmers described them as big racks full of custom chips built for one purpose: imitating, gate by gate, how a chip that does not yet exist will behave. Cadence's own description of its latest Palladium systems matches this: they run a representation of a chip before fabrication so engineers can debug the hardware and test software against it.

The speed difference explains why the emulator matters. Sohmers said simulating Positron's design in ordinary software, at what engineers call the register-transfer level (a code description of the chip's logic), runs at roughly 10 hertz, or about ten clock cycles per second. With some debugging features removed, a full-chip emulation on Palladium runs at around 500 kilohertz. That is about 50,000 times faster, the "several orders of magnitude" speedup Sohmers described.

What changed, he said, was the models. Sohmers credited GPT-6 Astra and Anthropic's Opus 5.5, saying the work would not have been possible three or six months earlier. He doubted the models had learned Palladium's documentation in training, partly because much of it is new with software updates. Yet the agents' reasoning traces and tool calls showed them reading the full documentation PDFs, condensing it into their own Markdown cheat sheets and then building the testing infrastructure and harnesses needed to use the machine.

For now, Sohmers said, the agents mainly write test programs, find cases where the design fails and write up reports. Other agents and humans then review those reports. He placed the turning point at roughly the beginning of September, when Astra came out, and called it a "step function" improvement over GPT-5.6, which still needed humans at different stages and could not close the loop itself. He said the setup brings Positron closer to a recursive self-improvement loop, in which AI helps build the hardware that runs AI.

The token bill

Asked how Positron's spending on AI compares with its spending on people, Sohmers said token spend, the money paid for the text models read and generate, had very quickly become the company's single largest line item outside manufacturing. When Labenz asked whether it was bigger than human salaries, Sohmers said it had recently surpassed them and then come back down.

He traced the curve. Six months earlier, token spend equaled about one employee's cost. By June it equaled several. At the peak, in the days after the new OpenAI model arrived and the team was pushing its limits, Positron was spending more than $100,000 a day on tokens. That eased in the following weeks as the team optimized its work and ran fewer experiments in parallel.

The "saving grace," Sohmers said, was Opus 5.5. On many of Positron's tasks, though not all, it did better than Astra at a quarter of the price.

The company never told anyone to cut back, he said, even while spending was growing exponentially, because it felt it was getting a good return. His own rule runs the other way from cost-cutting. He does not want anyone running a real software or hardware development task on anything less than the best model. A model that costs a tenth as much per token, he said, is "just not worth the expense."

No new ideas yet

Labenz asked whether the agents had come up with genuinely unexpected chip-design ideas, rather than following the rules in a manual. "I have not yet," Sohmers said, calling it "the most disappointing element of all."

He also said he was "somewhat proud" that the agents had not worked out some of Positron's cleverer design choices. They would flag a choice as a bad design decision until someone explained it, or until they ran tests and saw why the company does not build its systolic array the traditional way. A systolic array is a grid of simple processing units that pass data along in rhythm, a common layout for the matrix math at the heart of AI chips.

Sohmers was not sure that would hold for another six months. He did not expect an agent to have a better idea in a flash. Instead, he said, agents can iterate so quickly and run so many experiments on their own that they may reach the same conclusions his team did, or better ones, simply by trying far more design possibilities.

His comparison was OpenAI's recent work on the Navier–Stokes equations, which he described as sending 10,000 agents off for "millions of man-years" of effort until they found a solution. "It's brute force," he said. "It's expensive." But he called it a valid strategy that has only recently become possible. OpenAI's own account partly matches his description: the Navier–Stokes group involved roughly 10,000 agents working at once, and the reported result took 88 hours. The "millions of man-years" figure is Sohmers's framing; OpenAI's account does not give one.

Labenz closed the exchange by noting that not many people are left in the "AIs can't do what I do" category, and congratulated Sohmers on his place there "for as long as it lasts." Sohmers replied: "For a couple more weeks, at least."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

The Cognitive Revolution

AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Episode published This article draws on 40:50–48:15 (approximate times)

Article history

Updates to this article

Tags