Asked what comes after large language models, Alex told a caller on Moonshots with Peter Diamandis to separate two things people usually merge: the job of predicting the next piece of text, which he thinks has "effectively infinite longevity," and the transformer machinery doing it, which he says is already being swapped out part by part. Dave added his own forecast that the chips underneath will move to photonics within 18 months to two years.
Jakub Pachocki, OpenAI's chief scientist, published an essay saying no lab has solved alignment and monitoring well enough to keep scaling at full speed, and called for voluntary slowdowns until shared safety thresholds exist. On Moonshots, four panelists agreed the systems are extraordinary and disagreed with almost everything else in his argument.
OpenAI's GPT-6 Astra nearly saturates the interactive ARC-AGI-3 benchmark and leads Epoch AI's composite capability index, yet sits third on Artificial Analysis's suite, behind Claude Fable 5.1 and Muse Spark. On Moonshots EP #286, the panel works through what each ruler measures — and argues that Astra's real target was doing tasks with fewer output tokens, so a model can drive a desktop at conversational speed.
OpenAI has proposed ending the agreement that supplies its models to Cursor, now owned by SpaceX, on 12 November. On the Moonshots panel, one guest read the move as OpenAI betting on its own enterprise stack; another argued the real prize is reasoning traces — the working a model shows while solving a problem. Both explanations lead to the same awkward conclusion: Elon Musk and Anthropic now need each other.
Asked about allegations that Chinese labs extracted capabilities from Claude, Alvin Graylin argued that access to another model’s answers cannot explain every engineering advance. The Moonshots exchange turned on three distinctions: legitimate distillation versus prohibited extraction, query bills versus development costs, and learning from outputs versus improving the machinery behind them.
BDH-CQ’s authors report solving 118 of 400 public ARC-AGI-1 tasks at an estimated inference cost of $0.00070 per task, with up to two candidate answers. On Moonshots, Emad Mostaque welcomed architectural experimentation; panelist Alex questioned whether this design offered progress beyond a specialized benchmark.
xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.