16 September 2026
Heard In AI

Grok 4.6 closes the gap—and the panel asks what would take it ahead

xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

Two people had texted the host about Grok 4.6 before the recording. Introducing the release on Moonshots, he called it “a banger.” Alex, invited to comment first, offered applause—with a question about how far the recipe could go.

The released model was Grok 4.6, announced on August 12, 2026. Grok 4.7 reaching first place, as the episode’s title suggests, was still a prediction in the conversation. The panel described its arrival within two weeks as a rumor.

The Grok 4.6 announcement emphasizes long-running agent work. An AI agent does more than answer a question: it carries out a task through successive steps, using tools along the way. Here the promise is a system that can research, code, analyze or turn a broad product idea into a working first version. The company reports more self-testing and verification during extended tasks.

Grok 4.6 launched through Cursor, the AI coding environment, Grok Build and an API through which developers can connect it to their own software. Pricing starts at $2 per million input tokens and $6 per million output tokens. Tokens are the small pieces of text a model processes; input is what it receives, and output is what it generates. A faster variant costs twice as much.

What the number 61 means

The company’s launch table gives Grok 4.6 High a score of 61 on the Artificial Analysis Intelligence Index, alongside GPT-5.6 Sol Max at 61 and Fable 5 Max at 62. That places the named Grok configuration close to the leaders in that comparison, not alone at the top.

The index combines multiple evaluations into a weighted score. Its tasks and grading can change, so it is not a fixed exam with an unchanging scale. The launch announcement describes a nine-benchmark composite; Artificial Analysis’s v4.3 methodology describes ten evaluations across agents, coding, scientific reasoning and general capabilities. Those descriptions should not be treated as interchangeable without establishing which version produced the launch scores.

The v4.3 tasks include producing workplace deliverables, completing software workflows, writing scientific code and answering questions grounded in long documents. Depending on the task, grading uses executable tests, completion checks, answer comparisons or rubrics. Most evaluations measure first-attempt success, averaged over repeats where applicable. The suite is primarily English-language and text-based; visual, speech and multilingual abilities are measured separately.

So the launch comparison concerns particular model configurations on a particular collection of tests—not every capability a user might need.

The reasoning-trace argument

Alex opened his account of the training strategy with a caveat: “I lack insider information.” His reading was that Grok 4.6 was “essentially the next version of cursor.”

A reasoning trace is the intermediate working a model generates while solving a problem, rather than just its final answer. Alex argued that developers’ interactions with models through Cursor could supply valuable examples of that working, including traces from Claude and competing systems.

In his account, xAI’s acquisition of Cursor—which he understood was still being completed—and its licensing of Cursor’s trace data offered a shortcut toward leading-model performance. He compared it with allegations that Chinese labs were using Western models’ reasoning traces to improve their own systems. Both the Cursor-specific account and the comparison were his outside interpretation, not an independently established account of the training data.

Post-training means refining a model after its initial broad training, using examples and feedback to improve how it performs tasks. Teaching one model from another model’s outputs is a form of distillation. Alex’s argument was that worked examples from strong models can help a challenger reproduce skills it previously lacked, but do not by themselves supply a route beyond the systems that generated them.

“They won’t get you past the frontier,” he said, calling the approach a “one-trick pony” for nearly catching up. “But it’s a heck of a one-trick pony.”

The company’s documented training account is more specific about methods than data provenance. It describes supplemental training, regenerated examples of task-solving sequences, and reinforcement learning—learning from feedback or rewards—across coding, knowledge work, GPU software optimization, web development and computer-aided design. That account does not establish Alex’s broader description of which developers’ traces were used.

More chips, harder coordination

Alex’s optimistic case for Elon Musk’s strategy paired the post-training shortcut with access to NVIDIA GPUs, the processors used to train and run AI models. Learning from strong examples could close the gap; computing capacity could give the company room to attempt the next step.

The panel also discussed much larger models. Its account put Grok 4.5 at 1.5 trillion parameters, subsequently post-trained into 4.6, with 4.7 rising to 2 trillion and later models reaching 6 trillion and then 10 trillion. Parameters are the numerical settings learned during training—a measure of model size, not a performance score. Those figures were claims made in the discussion, rather than specifications established by the supplied release announcement.

Emad Mostaque’s explanation of the scaling problem centered on coordination. Writing training code, he argued, was less difficult than “trying to get 100,000 GPUs to do anything constructively together.”

Training a large model across many processors means their work must stay coordinated. Adding chips therefore creates a systems-engineering problem as well as adding computing power. Mostaque described fresh difficulties at successive scales and the time needed to get new hardware working effectively while meeting the intended price points.

That was his account of the bottleneck, not a confirmed explanation for the release delays the panel discussed. Alex also recalled Musk’s ambition to start a new pre-training run approximately monthly. He treated that as an unusually aggressive plan, not an established production schedule.

From a smarter model to a persistent teammate

The price brought the discussion back to the product. One panelist read the release as a move toward “persistent AI teammates”: systems that stay with work rather than simply produce a clever answer and stop. Lower-priced reasoning matters in that design because an extended task can involve many rounds of planning, tool use, testing and revision.

The panel expected Grok 4.7 to introduce another ingredient: SpaceX physics and engineering knowledge. That was an expectation about a future model, as was the claim that it would surpass Opus—not a result demonstrated by Grok 4.6.

For the model already released, the company’s stated direction is concrete: take a product idea through implementation, test the work, and keep going across the steps needed to produce a usable first version. The panel’s question was whether the next gains would come from better examples, larger training runs—or making that whole sequence work reliably at the price on offer.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Better data beat better architecture — but the panel split on its shelf life

A Moonshots panel unpacks Dwarkesh Patel and Jerry Han's experiment, which found that improvements in training data delivered a 12-fold compute-efficiency gain between 2019 and 2025 against 3.7-fold for architectures and training recipes — at small scale, on easy benchmarks. The panel then splits over whether a company's proprietary data is a durable advantage, with a $32 billion data-subsidiary valuation on one side and the fate of BloombergGPT on the other.

7 min read

DeepSeek's memory diet challenges what a data center needs to buy

On Moonshots #288, a 4 a.m. chart about DeepSeek's new V4.1-Flash model sent the panel from cache statistics to the shopping list for an AI data center. DeepSeek says the model's lookup memory needs a quarter of the expensive high-bandwidth memory and an eighth of the SSD cache storage of its previous generation. The panel's argument was about what that does to a buildout in which, by one panelist's estimate, 40% of American capital spending goes to that one component.

7 min read

Huang says AGI has arrived; OpenAI's 3.1 figure answers a narrower question

Nvidia's chief executive declared AGI achieved on September 6 while announcing more GPU capacity, and the Moonshots panel split between calling the label meaningless and calling the underlying capability the most important moment in history. A second claim on the same show — that OpenAI's agents now do 3.1 days of research work per human day — comes from an internal report that measures how long agents ran, not how much research they finished.

6 min read

Astra tops one leaderboard and trails another — the panel reads it as a computer-use model

OpenAI's GPT-6 Astra nearly saturates the interactive ARC-AGI-3 benchmark and leads Epoch AI's composite capability index, yet sits third on Artificial Analysis's suite, behind Claude Fable 5.1 and Muse Spark. On Moonshots EP #286, the panel works through what each ruler measures — and argues that Astra's real target was doing tasks with fewer output tokens, so a model can drive a desktop at conversational speed.

9 min read