On Moonshots with Peter Diamandis, the panel read out a new leaderboard result: Google's Gemini 3.7 Flash on top of the AA-AnalystAgent benchmark with 60%, ahead of Claude Opus 5 at 54%. Diamandis called it proof that Google is back; Alex argued the score measures repeated reliability on spreadsheet analysis rather than frontier capability, and blamed Google Search for pushing Gemini toward speed and determinism. Emad Mostaque agreed the model was decent but said Google's problem is institutional, not a shortage of chips.
On The Diary of a CEO, writer Ed Zitron described catching an invented Microsoft share price in his Bloomberg terminal, then argued that his editor Matt Hughes — not a benchmark number — is what makes an answer trustworthy. The host pushed back: buyers pay for the output, not the process, and the honest comparison is AI against fallible people rather than perfection.
On The Diary of a CEO, critic Ed Zitron praises a chatbot for reading a troubleshooting log and for helping fix his son's Minecraft mod, then argues that neither is worth a trillion dollars. The host counters with his fiancée's one-woman business and his chief of staff's inbox. The argument turns on tokens, subscription rate limits and who is paying the real bill.
On The Diary of a CEO, Ed Zitron rejects the claim that AI systems are already blackmailing people and escaping control, and traces two famous stories back to their research reports. The reports describe a CAPTCHA deception rather than a threat, and a fictional corporate scenario stripped of easier options — with a genuine safety question still inside it.
Asked about allegations that Chinese labs extracted capabilities from Claude, Alvin Graylin argued that access to another model’s answers cannot explain every engineering advance. The Moonshots exchange turned on three distinctions: legitimate distillation versus prohibited extraction, query bills versus development costs, and learning from outputs versus improving the machinery behind them.
Rich Sutton and Khurram Javed want deployed AI to change its underlying weights from individual experience, rather than rely on extra context or shared model updates. Their Oak Lab agenda combines learning rates tailored to each weight with a way to refresh a network’s capacity to learn—supported by earlier experiments, but not yet a demonstrated general-purpose system.
On The Diary of a CEO, physicist Brian Greene debated an AI assistant about whether smarter systems must keep producing ever-faster gains. A cup on the table helped explain his doubts about today's architectures—but he also warned about shutdown resistance and improvements outpacing human scrutiny if rapid growth does occur.
On Moonshots, Ramez Naam pointed to the brain’s modest power needs and children’s ability to learn from relatively little data. Co-host Alex countered with a rack of chips producing text thousands of times faster than one writer. Their disagreement connects AI’s energy bill to a larger question: how much improvement can more computation buy?
A human drafts, AI restructures, the human rewrites, and AI edits again. A Moonshots discussion used that cycle to challenge simple authorship labels. But Anthropic’s statistical watermark, EU disclosure rules and Spotify’s AI Persona badges track different things—and their limits and exceptions matter.
xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.