16 September 2026
Heard In AI

Area

Models & Research

Reporting and discussion about Models & Research, with links to the original sources.

Tags

Language models 13AI agents 11AI benchmarks 10Training data 9Scaling laws 8Emad Mostaque 7Inference costs 7Agent memory 6Anthropic 6Model distillation 6Multi-agent systems 5Open-weight models 5OpenAI 5Reasoning models 5AGI 4AI business models 4AI coding 4Human oversight 4Multimodal AI 4Recursive self-improvement 4Reinforcement learning 4Robotics 4Salim Ismail 4Superintelligence 4Transformers 4World models 4Agent harnesses 3AI mathematics 3Automated AI research 3Context engineering 3Continual learning 3Cursor 3Enterprise AI 3Google DeepMind 3GPT-6 Astra 3GPUs 3In-context learning 3Meta 3Synthetic data 3The Bitter Lesson 3US–China AI competition 3AI alignment 2AI chips 2AI pricing 2AI productivity 2AI risk 2AI scientific discovery 2AI startups 2AI video 2Claude 2Context windows 2Dario Amodei 2Data privacy 2Elon Musk 2Formal verification 2Khurram Javed 2Rich Sutton 2Sam Altman 2Simulation-to-real transfer 2Test-time compute 2The Alberta Plan 2Tool use 2Agent evaluation 1AI assistants 1AI control 1AI drug discovery 1AI in finance 1AI in law 1AI interpretability 1AI search 1AI slop 1AI-generated media 1AlphaGo 1AlphaZero 1Autonomous vehicles 1Catastrophic forgetting 1Codex 1Computer use 1Continual backpropagation 1Energy demand 1Gemini 1Google 1Hallucinations 1Jakub Pachocki 1Jensen Huang 1Local AI 1Meta-learning 1Navier–Stokes equations 1Neural plasticity 1NVIDIA 1Oak Lab 1Open-source AI 1Organizational design 1Qwen 1Research attribution 1The Big World Hypothesis 1Vibe coding 1World Labs 1

Aaron Levie's question for AI memory: what belongs in the weights?

On Training Data, Box CEO Aaron Levie was asked where enterprise AI memory is heading — retrieval, or models whose weights absorb a company's knowledge. His answer started with a lawyer who can see five matters and whose access changes daily, and ended with a wish for a rubric deciding what gets baked in and what stays a lookup.

6 min read

The prediction task may outlive the transformer, a Moonshots panel argues

Asked what comes after large language models, Alex told a caller on Moonshots with Peter Diamandis to separate two things people usually merge: the job of predicting the next piece of text, which he thinks has "effectively infinite longevity," and the transformer machinery doing it, which he says is already being swapped out part by part. Dave added his own forecast that the chips underneath will move to photonics within 18 months to two years.

6 min read

Better data beat better architecture — but the panel split on its shelf life

A Moonshots panel unpacks Dwarkesh Patel and Jerry Han's experiment, which found that improvements in training data delivered a 12-fold compute-efficiency gain between 2019 and 2025 against 3.7-fold for architectures and training recipes — at small scale, on easy benchmarks. The panel then splits over whether a company's proprietary data is a durable advantage, with a $32 billion data-subsidiary valuation on one side and the fate of BloombergGPT on the other.

7 min read

Huang says AGI has arrived; OpenAI's 3.1 figure answers a narrower question

Nvidia's chief executive declared AGI achieved on September 6 while announcing more GPU capacity, and the Moonshots panel split between calling the label meaningless and calling the underlying capability the most important moment in history. A second claim on the same show — that OpenAI's agents now do 3.1 days of research work per human day — comes from an internal report that measures how long agents ran, not how much research they finished.

6 min read

What OpenAI's 10,000 agents actually proved about fluid flow

OpenAI said on 8 September that an internal model, running roughly 10,000 agents for 88 hours, produced a forced blowup construction for the Navier–Stokes equations and a machine-checked proof of it. On Moonshots with Peter Diamandis, the panel worked through what the result is — a statement about idealized fluids, not a device — what it cost, and why the credit for it was contested within hours.

7 min read

World Labs' Atlas rebuilds a place from photos, and imagines the rest

World Labs released Atlas, a model that generates video along a camera path the user designs and rebuilds scenes from a handful of photographs. Its own garden example shows the seam: one photo leaves the surrounding buildings invented, while more photos pin them down. On Moonshots, the panel worked through what Gaussian splats are and why the approach might matter for robots, games and planning a vacation.

6 min read

How agent teams turned Fermat's proof into 13 million checked lines

On Moonshots with Peter Diamandis, a panelist interrupted an argument about AI regulation to read a headline off his feed: Anthropic had formalized Fermat's Last Theorem. Anthropic's report describes dozens of agents working eleven days, about six billion output tokens and 30,300 intermediate theorems — plus a piece of bookkeeping software that stopped runs from losing track of their own work. The panel's takeaway was about how to narrow enormous machine output into one result you can build on.

5 min read

Astra tops one leaderboard and trails another — the panel reads it as a computer-use model

OpenAI's GPT-6 Astra nearly saturates the interactive ARC-AGI-3 benchmark and leads Epoch AI's composite capability index, yet sits third on Artificial Analysis's suite, behind Claude Fable 5.1 and Muse Spark. On Moonshots EP #286, the panel works through what each ruler measures — and argues that Astra's real target was doing tasks with fewer output tokens, so a model can drive a desktop at conversational speed.

9 min read

Altman says AGI by year-end; the panel wants agents that stop forgetting

A TIME report has Sam Altman expecting an internal system he would call AGI within four months, and OpenAI's chief scientist saying its unreleased Astra model has met an internal benchmark for an automated research intern. On the Moonshots panel, the label mattered less than a practical test: whether the next model can finally keep hold of what it has learned over a long job, instead of handing a summary to a successor and starting again.

6 min read

A 100x claim lands, and the panel hits a harder question: coordinating 10,000 agents

On Moonshots, the panel revisits Elon Musk's January prediction that models were "off by two orders of magnitude" in intelligence per gigabyte, after Tim Sweeney tweeted that it had come true and Musk replied that specialist AIs add another 100x. Dave calls 100x a lower bound and asks what anyone would actually do with 10,000 brilliant agents; Emad Mostaque describes running specialized agent teams, while Alex argues Musk's "specialist models" are really sparsification inside generalist models.

6 min read

Why a Moonshots panel thinks China's AI tokens go to video and America's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

6 min read

What Gemini 3.7 Flash's analyst benchmark win actually measures

On Moonshots with Peter Diamandis, the panel read out a new leaderboard result: Google's Gemini 3.7 Flash on top of the AA-AnalystAgent benchmark with 60%, ahead of Claude Opus 5 at 54%. Diamandis called it proof that Google is back; Alex argued the score measures repeated reliability on spreadsheet analysis rather than frontier capability, and blamed Google Search for pushing Gemini toward speed and determinism. Emad Mostaque agreed the model was decent but said Google's problem is institutional, not a shortage of chips.

7 min read