Reflection AI has announced Beam, an open-weight model that the company says does the same work as its rivals with far less computation. On the October 7, 2026 episode of Moonshots with Peter Diamandis, recorded the day before, the panel spent less time on the model than on the question it raised: can an American lab compete in open models, and who would pay for it?
An open-weight model is one whose trained parameters, the billions of numbers that make up the model, can be downloaded and run by anyone. That differs from services such as Claude or ChatGPT, which customers can only reach through the company's own servers. In the panel's discussion, Chinese labs were treated as the current leaders in open-weight models.
What Reflection claims
Peter Diamandis, the founder of XPRIZE and Singularity University, who hosts the show, introduced Beam as "a new American contender." He described Reflection as an NVIDIA-backed lab founded in 2024 by former DeepMind researchers Misha Laskin and Ioannis Antonoglou.
Reflection's announcement, dated October 5, describes Beam as a "sparse mixture-of-experts" model. It has 501 billion parameters in total, but only 23 billion are active for each word-piece (token) it produces. The model sends each piece of work to a small set of specialized sub-networks instead of the whole network. Diamandis put it this way: "it thinks like a big model, but it runs like a much smaller model." Reflection says the model is aimed at coding, reasoning and agentic work, in which AI carries out multi-step tasks.
Diamandis summarized the headline claim as three to four times more efficient than China's GLM 5.2 and more than four times more efficient than the leading Western open models. "In plain English, it finishes tasks faster and cheaper," he said.
The efficiency claim depends on how Reflection measures it. The company estimates the computation needed to generate an answer as twice the number of active parameters multiplied by the number of tokens the model generates, including its reasoning steps. By Reflection's own description, that approximation leaves out the work of reading the prompt, costs that grow with longer contexts, and serving overhead. It is an estimate of computation, not of what it actually costs to run the model as a service. Reflection also says that some models, including China's Kimi K3, are still ahead on raw capability. When Beam was announced, it was in final evaluations with early-access registration. The weights, a model card and a technical report were planned for later in October.
Diamandis also cited Reflection's training figures: 10,500 NVIDIA chips, in what he said Reflection called the largest publicly documented reinforcement-learning run, with "no sign of a plateau." Reinforcement learning (RL) improves a model by rewarding good attempts at tasks. The announcement describes an RL campaign of more than 100 million attempts on 10,500 NVIDIA GB300 GPUs over four weeks, after pretraining on 23.8 trillion tokens.
Fewer tokens is not the same as cheaper
Alexander Wissner-Gross, a computer scientist and founder of Reified, said the independent benchmarking firm Artificial Analysis had published a preliminary analysis only hours before the recording. It found that Beam is likely one of the most token-efficient open models it has seen at that level of intelligence. He welcomed the focus but said token efficiency was "not necessarily the best metric." He would rather see cost efficiency. A model can use fewer tokens by doing more computation for each one, he said, for example with very deep or "looping" transformers. Transformers are the network design behind today's language models.
Artificial Analysis's published methodology keeps these measures separate. Its cost per task combines actual input, cached-input and output token use with each provider's prices. It reports response speed separately as well. Two models that produce similar numbers of tokens can therefore cost quite different amounts.
Wissner-Gross's bigger complaint was about capability. Right now, he said, American models are only "aspiring to just be on the cost or token efficiency frontier" rather than pushing what models can do. Without that, he expects that for the next few years Anthropic and OpenAI will lead on capabilities while everyone else optimizes for performance or cost. He called that "not the worst of all possible worlds, but not the best."
Mostaque: not aggressive enough
Diamandis asked Emad Mostaque, founder of Intelligent Internet and a longtime advocate of open models, whether America could win the open-weight race. "Of course they can," Mostaque said. He contrasted the 10,000 B300 chips used for Beam with the "few thousand hoppers," an older NVIDIA chip generation, that he said Chinese labs have. "They're just not going aggressive enough."
Mostaque called Beam "a good first try." In his reading, though, it underperforms Qwen 3.8 Next, a model a quarter of its size, as well as a DeepSeek Flash model. He also noted that Reflection compared Beam with GLM 5.2 rather than the newer 5.3. Reflection had raised $5 billion and this was its first release, he said, and he saw an "orders of magnitude difference" in price and capital raised, with money going into scale while Chinese labs work under the "necessity" of having less.
His advice had three parts. First, go "all in on the edge AI first": models that run on people's own devices, so that America has agents on everyone's devices that are American and work for Americans. Second, to reach the capability frontier, use distillation, in which one model is trained on another model's outputs. He named Kimi K3 as one source, calling it "perfectly legal payback." He also suggested working with Google to become the open equivalent of its models. Third, focus on data. With the same code and parameters as a stronger Chinese model, he said, the difference comes down to the training dataset, so Reflection could build datasets from Chinese models, work with US companies, and then try to win on scale and inference, the cost of running the model.
A gap in the market
Dave Blundin, founder and general partner of Link Ventures, saw demand. He described a US company that has decided it needs AI and doesn't want to "subscribe to Anthropic for the rest of my life." The only open alternatives, he said, are Chinese, and a bank or insurer may not know whether it is even allowed to use one. "You step in with an American product that's reasonably good. It's going to sell like crazy," Blundin said. He pointed to Reflection's valuation of about $25 billion and said he did not know whether it would succeed. But it was filling "a really, really wide open gap in the market, like a trustworthy U.S.-based open source platform that you can start from."
The conversation also turned to a model that companies could run on the hardware they already have, since many cannot get NVIDIA's newest chips. Blundin said a model distilled enough to run on a Mac "would be like a dream," though technically harder, and told Mostaque that if he built that company, "I'll invest in it like tomorrow." Salim Ismail, founder of Open ExO, guessed that 20 such companies were already working in stealth.
The business-model problem
Wissner-Gross returned to the money. His understanding, he said, is that Reflection is partly financed by NVIDIA and buys computing power from Elon Musk's Colossus 2 supercluster and from Nebius. Chinese open-weight labs, he said, earn money through value-added services contracts and government contracts. Some are also moving into hardware, selling custom memory and chips designed around their own models. An American lab could copy that approach through reseller or embedded-device deals, but "it's really tricky." If American open-weight labs can solve the business-model problem, he said, they can solve the American open-weight problem. "But we haven't solved that in the West yet."
Diamandis asked whether the big closed US labs would release open-weight models. "No," Blundin said. Wissner-Gross said they already have, pointing to Google's Gemma, "but it's not competitive."
The two disagreed on the reason. For Blundin, it is liability. In his account, the White House had said developers are fully liable for whatever they release. A Google, Anthropic or OpenAI would therefore not release an open model when "anyone could do anything with that." For a large company, he said, the upside is "so small compared to the risk of a massive lawsuit," while for a startup releasing models is the business.
He gave an example from his own experience. A startup he was involved with years ago ran into an old California law he called the "Fred Astaire law," written with billboards in mind. A class-action lawyer counted each internet ad impression as a separate instance, arrived at 400 million, and claimed liability of $200 billion against a 10-person company. Blundin said they paid the lawyers a sum of money and the claim went away. Laws like that, he said, "predate the internet, let alone AI."
Wissner-Gross put it down to economics. He pointed to a practice called obliteration, a blend of "ablation" and "obliteration." People use freely available tools to post-train an open-weight model so that its guardrails are removed. The model then refuses less often, and he said its capabilities improve. For Wissner-Gross, though, the main reason closed labs hold back is simpler: "I'm not even sure if it's about safety or alignment or liability. It's just like, what's the point business-wise?"
Where the money is
The discussion turned to who already profits from open models. Together AI, Modal and Baseten were cited as running at billion-dollar revenue rates by hosting open models, with the remark that it would be good if they ran American models rather than Chinese ones. NVIDIA's spending on open-source models was put at $20 billion. Another figure raised was that Reflection is paying $1.5 billion a year for computing, while Chinese labs do the same on a quarter of that. Wissner-Gross said those hosting companies are "in the infra business," not the model business, where "you can make enormous amounts of money, forget about profit, but revenue."
Blundin's view was broader. "Every corporation is about to go into panic mode saying, what's my AI strategy?" he said. Whether a company sells models, computing, data centers or consulting, "as long as you're in the room, when they have that panic moment, you're going to sell and you're going to succeed." The panel's example was Mistral, the French lab, described as having reached a billion-dollar revenue run rate by being in the room with European companies and selling them long-term services contracts, and as having just released a model comparable to Reflection's.