16 September 2026
Heard In AI

Tag

Language models

Articles about Language models from podcasts, articles and papers, with links to the original sources.

What Gemini 3.7 Flash's analyst benchmark win actually measures

On Moonshots with Peter Diamandis, the panel read out a new leaderboard result: Google's Gemini 3.7 Flash on top of the AA-AnalystAgent benchmark with 60%, ahead of Claude Opus 5 at 54%. Diamandis called it proof that Google is back; Alex argued the score measures repeated reliability on spreadsheet analysis rather than frontier capability, and blamed Google Search for pushing Gemini toward speed and determinism. Emad Mostaque agreed the model was decent but said Google's problem is institutional, not a shortage of chips.

7 min read

Why Ed Zitron trusts his editor more than a hallucination score

On The Diary of a CEO, writer Ed Zitron described catching an invented Microsoft share price in his Bloomberg terminal, then argued that his editor Matt Hughes — not a benchmark number — is what makes an answer trustworthy. The host pushed back: buyers pay for the output, not the process, and the honest comparison is AI against fallible people rather than perfection.

7 min read

Would she pay the real price? Zitron's test for AI adoption

On The Diary of a CEO, critic Ed Zitron praises a chatbot for reading a troubleshooting log and for helping fix his son's Minecraft mod, then argues that neither is worth a trillion dollars. The host counters with his fiancée's one-woman business and his chief of staff's inbox. The argument turns on tokens, subscription rate limits and who is paying the real bill.

7 min read

What the AI blackmail experiments actually tested

On The Diary of a CEO, Ed Zitron rejects the claim that AI systems are already blackmailing people and escaping control, and traces two famous stories back to their research reports. The reports describe a CAPTCHA deception rather than a threat, and a fictional corporate scenario stripped of easier options — with a genuine safety question still inside it.

5 min read

Why Graylin says distillation cannot explain all of China’s AI gains

Asked about allegations that Chinese labs extracted capabilities from Claude, Alvin Graylin argued that access to another model’s answers cannot explain every engineering advance. The Moonshots exchange turned on three distinctions: legitimate distillation versus prohibited extraction, query bills versus development costs, and learning from outputs versus improving the machinery behind them.

6 min read

Oak Lab wants AI that keeps learning from you

Rich Sutton and Khurram Javed want deployed AI to change its underlying weights from individual experience, rather than rely on extra context or shared model updates. Their Oak Lab agenda combines learning rates tailored to each weight with a way to refresh a network’s capacity to learn—supported by earlier experiments, but not yet a demonstrated general-purpose system.

7 min read

Is the brain more energy-efficient than AI? It depends what you count

On Moonshots, Ramez Naam pointed to the brain’s modest power needs and children’s ability to learn from relatively little data. Co-host Alex countered with a rack of chips producing text thousands of times faster than one writer. Their disagreement connects AI’s energy bill to a larger question: how much improvement can more computation buy?

6 min read

Grok 4.6 closes the gap—and the panel asks what would take it ahead

xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.

6 min read