Keerthana Gopalakrishnan, research lead for Gemini Robotics at Google DeepMind, has a simple test for whether a robot "brain" deserves the name. Train it on one machine, then move it to another: a humanoid, or "some other robot that I just bring in." If it is "completely helpless in that setting," she asked, "is that a generic brain?"
By that test, she said on The Cognitive Revolution, in an episode published October 3, 2026, robotics is "still GPT-2." GPT-2 and GPT-3 were early OpenAI language models. GPT-3 became known for few-shot learning. Its 2020 paper showed that a model could pick up a new task from a handful of examples placed in its prompt, with no retraining. Gopalakrishnan's point was that robots have not yet reached a comparable stage.
The conversation came after DeepMind's July 30, 2026 release of Gemini Robotics 2. The suite includes a vision-language-action model, or VLA, which turns camera images and instructions into robot movements.
Why half the time is not enough
The discussion started with a measure taken from software agents. METR, a group that evaluates AI systems, charts how long a task can be (measured by how long it takes a human) before a model's success rate drops to 50% or 80%. The question put to Gopalakrishnan was that a 50% success rate may be tolerable on a laptop, where a missed detail is a nuisance. A robot that drops something breaks it. So which robot tasks are reliable 99% of the time or more, and which sit nearer 80% or 50%?
Gopalakrishnan said a lot of pick-and-place work, which means moving objects from one spot to another, is "close to deployment." Then she changed the frame. Generalization and mastery, she said, are "somewhat orthogonal axes." A team can build very narrow AI that is excellent at one or two things. But then "the cost of doing the nth thing becomes exactly the same as the cost of doing the first or the second thing."
A broad baseline changes that cost. Gopalakrishnan described working on generalization as "lifting the wave for all the boats": general common sense and understanding make it easier to reach mastery on every task. She pointed to language models. People once built specialized models for single jobs, she said, but GPT and Gemini now handle a wide range of tasks, "and then from those general baselines, then you can climb really high." She said she thinks "the same thesis might hold true for robotics as well."
Showing a robot once
Several robotics companies have posted videos claiming few-shot or even one-shot learning, where a person does a task once and the robot copies it. Gopalakrishnan was asked what deploying a robot will really take: a hundred demonstrations, ten, or one?
"I think it's going to be a spectrum," she said. Teaching a robot in context, meaning by showing rather than retraining, shortens the time it takes to pick up a task and to deploy it, and she called that "a very exciting development." But she described it as one more way of prompting a model. Telling an older model to pick up an object it has never seen is already a request for general behavior through language. Drawing a circle around the item to be handled is prompting with an image. A demonstration is "prompting with video in some sense."
The open question, she said, is how far these models can generalize from such prompts. They are prone to memorization: "if you really show it exactly, this is what to do, they will copy it." Two tests matter. First, can the robot adapt when the scene changes? Second, how hard is the task? Models have plenty of data for pick-and-place. "But can you show a robot to tie a trash bag?" she asked. "I think it's still fairly early."
Two conditions for a GPT-3 moment
Asked to place robotics on a scale of one to six GPTs, Gopalakrishnan said GPT-2 and named two things that would be needed for GPT-3. Few-shot learning has to work well across many different tasks. And robotics, she said, "weirdly is still very subject to cross embodiment": whether a model trained on one body works on another.
Language models never had that problem, she said. "My phone or your phone, my computer, Mac, Linux, it doesn't matter where you run it." Robot models, by contrast, "are very subject to which robots that you act on," and much work remains to build "very generic brains that can discount those factors out."
Part of the reason is data. Gopalakrishnan said VLAs still train on data collected from robots. DeepMind is building its VLA with partners she named as Apptronik, Agile Robots and Boston Dynamics. Apptronik's partnership dates back to the first Gemini Robotics release in March 2025. Cross-embodiment has taken "a large leap," she said. But taking a new body and zero-shotting it to many tasks, meaning succeeding with no examples from that body, "with very high reliability" is something for which she sees no strong precedent.
What works today is fast adaptation. Gopalakrishnan described a trusted-tester program for the Gemini Robotics On-Device model, which runs on the robot itself, through which outside labs use it with their own robots. She called one result "one of my favorite parts" of that release: with 200 examples, she said, the model can learn many different tasks on many different bodies, because it already has a general understanding of the physical world and of controlling different machines. Google's On-Device 2 page says the model adapts to completely new robots with fewer than 200 examples. The Gemini Robotics 2 announcement puts it more narrowly: it describes adapting to new bi-arm robot embodiments with a few hours of adaptation time, typically with fewer than 200 examples, even when shapes, sensors and degrees of freedom differ drastically. That announcement also shows one shared VLA checkpoint tested on an Apollo humanoid with two different kinds of hands and on a two-arm Franka setup with grippers.
Drones with arms
Asked about the most exotic bodies to come through the tester program, Gopalakrishnan said there were "surprisingly, not a lot." The weirdest she recalled was "people putting arms on drones." She still thinks unusual bodies can be handled: "if you have data, you can learn anything. I think that is the principle."
She described the range of robot bodies as something like a normal distribution. The common types cluster together: two-armed setups, hands, humanoids. "Everything weird is like a long tail."
The fried-egg problem
Could one of her robots fry an egg, plate it and bring it to the table? "It depends on whether you have data or not," Gopalakrishnan said. For a single job, there are algorithms that can make a robot very reliable. A general model that has never seen eggs being fried, set loose in a stranger's kitchen, is a different matter. She said models have made a lot of progress on visual generalization and are no longer thrown off by lighting or the layout of different homes. The harder part is task meaning, context and "the ability to make mistakes."
Frying leaves little room for error: "you can overcook an egg and no one would like that." Tasks that allow retries are much easier. In Lego assembly, she said, a mistake can be undone. "If you crack an egg and then it's on the floor," she added, the robot has no way back.
Why build humanoids at all
Gopalakrishnan acknowledged that many people ask why researchers bother with humanoids instead of simpler robot arms. Her answer was that whole fields of research stay out of sight until you aim at "the peak form factor." She said DeepMind is "probably the only lab here in North America" where humanoid intelligence can be studied across many bodies. The lab has several humanoids and is not building a brain for any one of them. The aim, she said, is to work on the intelligence while abstracting away the body a little.
Working with different humanoid bodies brings three problems to the front, she said. Human-robot interaction "becomes a much bigger thing." So does whole-body control, which means coordinating legs, torso and arms at once. And multi-finger dexterity becomes "a new factor." To be exposed to all three, she said, "you need to work on it."