Asked what comes after large language models, Alex told a caller on Moonshots with Peter Diamandis to separate two things people usually merge: the job of predicting the next piece of text, which he thinks has "effectively infinite longevity," and the transformer machinery doing it, which he says is already being swapped out part by part. Dave added his own forecast that the chips underneath will move to photonics within 18 months to two years.
A Moonshots panel unpacks Dwarkesh Patel and Jerry Han's experiment, which found that improvements in training data delivered a 12-fold compute-efficiency gain between 2019 and 2025 against 3.7-fold for architectures and training recipes — at small scale, on easy benchmarks. The panel then splits over whether a company's proprietary data is a durable advantage, with a $32 billion data-subsidiary valuation on one side and the fate of BloombergGPT on the other.
Nvidia's chief executive declared AGI achieved on September 6 while announcing more GPU capacity, and the Moonshots panel split between calling the label meaningless and calling the underlying capability the most important moment in history. A second claim on the same show — that OpenAI's agents now do 3.1 days of research work per human day — comes from an internal report that measures how long agents ran, not how much research they finished.
OpenAI's GPT-6 Astra nearly saturates the interactive ARC-AGI-3 benchmark and leads Epoch AI's composite capability index, yet sits third on Artificial Analysis's suite, behind Claude Fable 5.1 and Muse Spark. On Moonshots EP #286, the panel works through what each ruler measures — and argues that Astra's real target was doing tasks with fewer output tokens, so a model can drive a desktop at conversational speed.
On Moonshots, the panel revisits Elon Musk's January prediction that models were "off by two orders of magnitude" in intelligence per gigabyte, after Tim Sweeney tweeted that it had come true and Musk replied that specialist AIs add another 100x. Dave calls 100x a lower bound and asks what anyone would actually do with 10,000 brilliant agents; Emad Mostaque describes running specialized agent teams, while Alex argues Musk's "specialist models" are really sparsification inside generalist models.
On the Moonshots panel, Peter Diamandis runs a Grok Bot chief of staff called Skippy and Emad Mostaque runs 18 of them across his own machines, installing models and making art. Salim Ismail calls it the move from asking an AI to assigning work; Alex argues the messaging-app interface cannot possibly scale.
On The Diary of a CEO, physicist Brian Greene debated an AI assistant about whether smarter systems must keep producing ever-faster gains. A cup on the table helped explain his doubts about today's architectures—but he also warned about shutdown resistance and improvements outpacing human scrutiny if rapid growth does occur.
On Moonshots, Ramez Naam pointed to the brain’s modest power needs and children’s ability to learn from relatively little data. Co-host Alex countered with a rack of chips producing text thousands of times faster than one writer. Their disagreement connects AI’s energy bill to a larger question: how much improvement can more computation buy?
BDH-CQ’s authors report solving 118 of 400 public ARC-AGI-1 tasks at an estimated inference cost of $0.00070 per task, with up to two candidate answers. On Moonshots, Emad Mostaque welcomed architectural experimentation; panelist Alex questioned whether this design offered progress beyond a specialized benchmark.