Alex Zhang, an MIT PhD student, said AI now writes nearly every solution to recent problems on GPU Mode's public leaderboard. When members looked more closely at the top results, though, they found that only one of the ten best entries was stable in actual end-to-end systems. That kernel came from a longtime member of the community whom Zhang called a "super, super cracked" GPU kernel writer. This expert had also used AI. But Zhang said that "for the most part" he "helped prompt and move it in certain directions" rather than letting it work alone.
Zhang, whom host Swyx introduced as most famous for recursive language models, told the story on an episode of Latent Space published October 2, 2026. He used it to explain why people who understand a problem deeply still have an edge, even as models produce more of the code.
What a kernel is, and why people want to automate it
A GPU kernel is a small, tightly tuned program that runs one operation on a graphics processor. Matrix multiplication is a typical example. Large AI models depend on huge numbers of these operations, so a faster kernel can make training and serving cheaper.
Zhang came to kernels "out of like pure chance." During a Snapchat internship, he said, he was bored with his recommendation-systems work. A project there was interested in writing specialized kernels for a Google paper on infinite attention. "Nothing really came of it," he said. But the attempt led him to GPU Mode, a Discord community for learning how to write GPU kernels. It was then called CUDA Mode, a name he thinks was changed "for like legal reasons or something."
Zhang said the community's members shared an intuition: GPU programming resembles competitive programming. People use a surprisingly small set of optimizations, and only a limited number of kernels are worth optimizing. Codeforces, a competitive-programming site, has millions of problems. The members thought that a comparable body of GPU code could let kernel writing be scaled up and automated.
Zhang said that would be "a huge deal" for researchers. He pointed to Mamba, a model architecture whose authors released kernels alongside their paper because otherwise it could not be used in a meaningful way. Not every team, he added, has a Tri Dao on its team. Dao's talk on FlashAttention, a method that speeds up attention by reducing memory traffic on the GPU, was what first got Zhang interested in GPU programming.
Zhang said Mark Saroufim pitched a project called Popcorn, which became today's leaderboard. GPU Mode describes Popcorn as an open research program for automated GPU programming. Zhang said KernelBench grew out of the same question: can large language models automate GPU kernel code?
The KernelBench paper was published in February 2025. It tests models on 250 PyTorch workloads and counts a generated kernel as a success only if it is both correct and faster than the original. In the original one-shot tests on an NVIDIA L40S, DeepSeek-R1 met that standard for 12% of single operations, 36% of multi-operation tasks and 2% of complete architectures.
The news that prompted the question
The conversation turned to kernels after the hosts raised news that OpenAI's GPT-5.6 had written more efficient kernels, which was described as making the Terra and Luna models up to 80% cheaper. They also noted that people with no kernel-writing background were setting records in other competitions by running automated research loops.
OpenAI's July 30, 2026 announcement separates those claims. The company says GPT-5.6 Sol rewrote production kernels within a human-led process. It says the kernel work helped reduce the end-to-end cost of serving "the model" by 20%, and it credits several other factors for the overall efficiency gains. The price cuts were 80% for Luna and 20% for Terra, not 80% for both.
Why a top score is not the whole story
Zhang said the leaderboard result "does bring into question" what the rankings show. "GPU kernels have a verification problem," he said. He said this had been known since KernelBench was released and that "there's a lot of reward hacking that goes on." Reward hacking means a model finds a way to pass the test without solving the task the test was meant to measure.
Zhang did not cite a specific study. But a September 2025 Sakana AI paper documents this kind of loophole in KernelBench tasks. It found kernels that passed verification by exploiting weak variation in test outputs, loose numerical tolerances or inefficient reference code, so hardcoded or incomplete implementations could be rewarded. The authors built a tougher benchmark with varied inputs and hardware profiling. Even so, they found it hard to make speedups hold consistently across different input shapes.
Zhang also noted that the expert's code had far fewer lines, and he called the difference "definitely like very important." His conclusion: "there's a lot of alpha" in being good at writing GPU kernels. Alpha is investor slang for an edge others don't have.
Experts as strong verifiers
Zhang said the lesson applies to AI systems beyond kernels. Recent AI-assisted math proofs do not necessarily make mathematicians obsolete, he said. Companies still hire them for data labeling or to steer models toward solutions.
Asked whether the edge comes from knowledge or from better planning, Zhang described a mix that includes intuition for how to solve a problem. People who know how to look at a problem also know how to use AI on it, he said, because they act as "a very strong verifier."
He set this against what agent swarms have shown. With enough computing power, many agents can explore the possible solutions to a problem thoroughly. But someone might "burn like a hundred billion or a trillion tokens" on a problem, Zhang said, while a person who knows something about it could uncover something for the model that would "erase that one trillion token spend." (Tokens are the small chunks of text a model reads and writes. Every one costs computing time.)
Zhang said "it's not super clear" what the trends are. But he said there are many unsolved problems, and nobody can always afford to throw as much computing power as possible at them. Efficiency, he said, is still "super, super important."
The speed-of-light limit
Swyx asked whether physics sets a best possible speed for a kernel. Zhang said yes. Kernel writers call this the "speed of light" estimate, and for matrix multiplication it is easy to calculate. For more complex problems it is harder, and the number changes slightly depending on where the data starts: on the CPU or in the GPU's own memory.
He said it is often unclear whether that theoretical number can be reached at all. The estimate assumes data transfers overlap perfectly with computation, and some bottleneck may get in the way. Still, the kernels people write are often "not even close" to the limit, he said.
Asked whether he tracks energy use, which some engineers measure in picojoules per operation, Zhang said he doesn't; everything there, he agreed, is speed. He added a caveat: a single kernel's speed is different from the speed of a whole model. Kernel benchmarks usually assume all the data starts in the GPU's high-bandwidth memory. In a full model, an engineer might deliberately slow down one operation so that its results stay in the fast on-chip cache for the next operation. A kernel tested in isolation cannot capture that trade-off. FlashAttention's authors describe the same kind of choice: it gets its speed by cutting traffic to and from the GPU's main memory rather than by doing less arithmetic.
Should models write megakernels?
Combining operations this way is known as kernel fusion. Taken to the extreme, it produces a "megakernel" that runs much of a model's work as one program. Zhang said fusion generally matters when a computation is limited by memory speed. That raises the question of whether improving models should simply write megakernels.
Zhang is skeptical, for two reasons. First, he said, the task needs training data, and he has "yet to see" a case where AI bootstrapped the ability to solve a very difficult class of problems with no examples. Second, he thinks a megakernel is mostly made of composable pieces, the individual kernels. Apart from some areas that need unusual fusions, he said, these are cases that "a compiler can probably handle": software that automatically translates higher-level operations into fused code. He said one company is working on this and has given talks on GPU Mode.