On Training Data, Box CEO Aaron Levie describes how his customers actually pick models: a default for asking questions of their files, and hard-nosed accuracy evaluations for the high-volume extraction work where most tokens are spent. He endorses Decagon founder Jesse Zhang's argument that mature workflows migrate to open-weight models, and explains why the big labs' revenue and open-weight token volume can climb at the same time.
OpenAI's claimed Millennium Prize result used roughly 10,000 agents on a problem that was, as one entrepreneur on Moonshots put it, unusually easy to specify. The panel's argument: as the price of that kind of compute falls, the scarce skill becomes writing the target — and today's models, asked for ten ideas to cure cancer, produce a bad list.
On Moonshots, the panel picked apart a rental index showing H100 prices rising 22% in a single month to $3.28 per GPU-hour. Dave called it a reversal of a lifetime of chip depreciation; Emad Mostaque explained why better models make the same old Hopper worth more; and the warning for companies was that the compute they assume will be there later is already sold out.
On Moonshots #288, a 4 a.m. chart about DeepSeek's new V4.1-Flash model sent the panel from cache statistics to the shopping list for an AI data center. DeepSeek says the model's lookup memory needs a quarter of the expensive high-bandwidth memory and an eighth of the SSD cache storage of its previous generation. The panel's argument was about what that does to a buildout in which, by one panelist's estimate, 40% of American capital spending goes to that one component.
On the Moonshots panel, Alex argued that China's AI loyalty perks are the start of "universal basic tokens" — redistributed access to machine intelligence — while another panelist countered that cheaper tokens will mean bigger bills, not free ones. Emad Mostaque pushed past access to ownership, proposing 100 million publicly underwritten robots owned by the people, and Alex said that sounded like communism.
Anthropic's Fable 5.1 charges $0.25 per million tokens for cached reads, a quarter of the previous rate, which one Moonshots panelist read as an invitation to load an entire company's context into the model and keep it there. The panel connected that price to a wider scramble: with model leads lasting about a month, the labs are racing to convert them into customer workflows, partnerships and proprietary design data that a rival cannot copy.
On The Diary of a CEO, David Friedberg argued that open-weight AI models will stop the industry's value from pooling in two or three labs, and wagered that someone with no money today will build a billion-dollar company on a model they downloaded. His case runs through the Netscape era, the fight in Washington over Chinese models, and a proposal that data centers generate their own power and sit in ordinary retirement accounts.
Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.
On Moonshots with Peter Diamandis, Salim Ismail argued that a Mac Studio with 512GB of unified memory changes AI spending from a perpetual per-token bill into a capital asset, with law firms and mid-sized healthcare organizations as the likely buyers. Two other panelists agreed the machine was worth having and still called Apple's AI record a long-running software failure.
On The Diary of a CEO, critic Ed Zitron praises a chatbot for reading a troubleshooting log and for helping fix his son's Minecraft mod, then argues that neither is worth a trillion dollars. The host counters with his fiancée's one-woman business and his chief of staff's inbox. The argument turns on tokens, subscription rate limits and who is paying the real bill.
On Training Data, Parag Agrawal explains how his company Parallel entered web search without first building a giant index: it launched a search agent that crawled after a request arrived, replaced outsourced human data collection for insurance, sales and finance customers, and treated the index as a latency optimization to be grown later. He describes the agent-specific architecture behind it, the 200-millisecond Turbo mode Parallel announced in July, and a Google Cloud deal that puts Parallel Search beside Google Search as a grounding option.
Alvin Graylin argues that AI can become more useful while earning less for the companies financing its infrastructure. His warning centers on cheaper models and local computing weakening cloud revenues, just as NVIDIA proposes financing platforms intended to mobilize more than $500 billion of outside capital.