How many AI agents could the world's chips keep running at the same time? A new estimate from the research group Epoch AI answers in tens of millions for today's most capable models. If cheaper open models were used instead, the figure rises into the billions. On an episode of Moonshots with Peter Diamandis, recorded on October 6, 2026 and published the next day, the hosts and their guest used the number to argue about what a digital workforce of that size would mean. They also discussed the component that limits it: memory.
What Epoch counted
An AI agent is a model that carries out tasks on its own: it plans, uses software tools and works through steps without a person prompting each one. While an agent runs, it occupies part of a chip's high-bandwidth memory, or HBM, the specialized fast memory placed next to AI processors. That makes memory a practical way to count how many agents could run at once.
Peter Diamandis, the show's host and founder of XPRIZE, presented the study as a story that "reframes the labor debate." In Epoch's report, published October 2, 2026, researcher Jason Li starts with the HBM expected to ship from 2025 through 2027. He converts that supply into the equivalent number of Nvidia GB300 chips. He then estimates how many agents that hardware could serve, using serving benchmarks for open models and estimates for closed models inferred from API spending (the paid interfaces software uses to call a model). Across his hardware and serving assumptions, the result is about 30 to 170 million frontier-model agents running at the same time.
The estimate depends on several conditions. It assumes the hardware is fully deployed and devoted to these workloads. Delays in deployment and chips used for other purposes would lower the real number. Epoch's conclusion also considers the opposite of scarcity: demand for running AI models might not grow enough to use everything that gets built.
A separate scenario uses DeepSeek V4 Pro, an open-weight model whose files anyone can download and run. On the same pool of hardware, that model could support about 1.9 billion agents. Diamandis described this as "the working hour equivalent of 8 billion humans." The arithmetic holds. An agent that runs around the clock logs 168 hours a week, or 4.2 times a 40-hour work week, and 1.9 billion times 4.2 is about 8 billion. Epoch compares hours the same way for its main range: 30 to 170 million agents running continuously equal the weekly hours of roughly 140 to 720 million people working 40-hour weeks. The report says this is equal working time, not equal productive output.
From a billion to a trillion
Alexander Wissner-Gross, a computer scientist and founder of Reified, went beyond the report. Depending on how the strength of American closed models or Chinese open-weight models compares with that of human workers, he said, "we're either at or about to be, order of magnitude, a billion full-time equivalent humans that are AI agents." He then extended the trend: 10 billion "in a year or two," then 100 billion, then a trillion, "by the end of the decade, you extrapolate out."
"It doesn't take a rocket scientist to be a rocket scientist or to extrapolate straight lines," Wissner-Gross said. This is his own forecast, built on his own assumptions about how capable the models are. It is not Epoch's finding, because Epoch counts hours of possible operation, not human-equivalent ability.
Not competent enough yet
Diamandis asked guest Emad Mostaque, founder of Intelligent Internet, whether human labor is "cooked." Mostaque was more cautious about timing. "It's coming," he said, but "it's not quite there yet because even these models are not quite competent enough." He expects the next generation to be capable enough, and to be quantized, a technique that stores a model's numbers at lower precision so that it needs less memory.
Mostaque said DeepSeek V4 Pro needs 256 to 500 gigabytes of RAM. Within a year, he expects similar capability in a model the size of Alibaba's Qwen models that works on a MacBook, and "then everyone's got an agent." He said the internet company Cloudflare already sees more traffic and requests from automated agents than from humans. Once everyone has "an agent or 10," he expects agents to overtake humans in economic transactions "literally within a few years." In his picture, people start with an excellent personal assistant, then add a chief of staff, and then they are "marshalling entire teams."
The wrong unit of analysis
Salim Ismail, founder of Open ExO, said asking whether AI replaces a particular job is "the wrong unit of analysis." He tied the bigger shift to what the group calls the organizational singularity. "What happens when a company can summon 100,000 competent digital workers overnight?" he asked. A company could ask for 50,000 developers, 20,000 marketers and 5,000 legal experts for 48 hours "and then turn them off."
Ismail described a "second workforce" that "can be copied, works 24 hours a day, doesn't unionize, doesn't get sick, doesn't take coffee breaks, improves every quarter, and costs almost nothing." He argued that coordinating these workers becomes the main constraint. "The big challenge is not going to be the intelligence," he said. "It's going to be, what the hell do you ask them to do?"
Memory as the choke point
Dave Blundin, founder of Link Ventures, focused on the hardware underneath all of this. "It's amazing to me that the constraint to all of intelligence and therefore all of human progress is RAM," he said, referring to high-bandwidth memory. As evidence of how acute the shortage is, he pointed to investor interest in startups that work around it.
The first was Positron. Blundin said he was on stage with its founder, Mitesh Agrawal, the week before last. He described the company as 16 months old, valued at $5 billion and having raised almost a billion dollars because "they found a way to not use HB RAM to do some AI inference." Inference is the work of running a trained model to produce answers. Positron's September 10, 2026 announcement confirms $875 million in financing at a $5 billion valuation. That total is a $375 million Series C round plus a later tranche of up to $500 million. The announcement says the company's next-generation systems use commodity LPDDR5X memory, the type found in ordinary devices, to avoid depending on scarce HBM. Those chips are not yet in production: Positron plans to finish their design by the end of 2026 and begin production in the second half of 2027.
Blundin's second example was Alpaca, a startup he said was founded by an MIT student just starting his junior year. By Blundin's account, it has an eight-figure valuation and 20 to 30 employees, and it has found a way to ease the memory bottleneck "just a little bit."
Diamandis added that the memory maker SK Hynix is about to develop capabilities in the United States, and that Elon Musk has said RAM is the constraint on the future of AI. "Literally the entire constraint to progress in all fields is now tied up in one thing, just RAM," Blundin said. "Can we make more RAM?"