Keerthana Gopalakrishnan, research lead for Gemini Robotics at Google DeepMind, described a robot "brain" split into two parts. One model works out what should happen. A second model turns that plan into movement of a robot's whole body, "from fingertips to the feet." Speaking on The Cognitive Revolution, in an episode published October 3, 2026, she explained how the parts talk to each other, why only the reasoning half is open to developers so far, and why chaining the two models together creates new ways for a task to fail.
Three models, two jobs
Google DeepMind announced the suite on July 30, 2026. It has three parts, and Gopalakrishnan described each one.
Gemini Robotics ER 2 (ER stands for embodied reasoning) is what she likened to a "system 2" brain: the slower, deliberate thinker, a term borrowed from psychology's fast-and-slow model of the mind. It does "very generic reasoning." She said it is based on Google's Flash line of models, tuned toward robotics, and can also be asked general questions about images and video. Google's model card confirms that it is built on Gemini 3.5 Flash with extra training for embodied reasoning. It accepts text, images, video and audio, and it answers only in text, including structured instructions and tool calls. In other words, ER 2 never moves a motor itself. It decides what should happen and asks other software to do it.
Gemini Robotics 2 is a vision-language-action model, or VLA: a model that takes in camera images and an instruction and outputs robot movements. Gopalakrishnan described it as the model controlling "how the robot moves and what it should do given a certain intention from either a user or another" reasoning model.
Gemini Robotics On-Device 2 is, in her words, "a much smaller version" of the VLA that "fits on the robot's computer" rather than running in the cloud. Google says it is intended for situations where network delay or unreliable connections make cloud control a poor fit.
Access is split the same way. ER 2 can be used by anyone through the Gemini API and Google's AI Studio, while both action models are limited to partners and trusted testers.
A reasoning model anyone can call
Because ER 2 is open, the episode includes a simple demonstration of it. A coding agent connected the model to a simple 2D scene drawn in a web browser and gave it two tools, "grab at" and "place at." Each tool required the model to pick exact x-y coordinates from the image. The setup took only a few prompts and the model handled the job well. In one test the object was moved just before the grab, as a nod to well-known robot videos. The grab failed, the model received a new image, and it found the object again. Google's own orchestration tutorial follows the same pattern. It declares one tool for moving to coordinates and another for opening or closing a gripper. The developer's code runs each requested action and sends the result back to the model, round after round, until the block is in the bowl.
Without being asked, the coding agent also ran the same scenes on the base Gemini 3.5 Flash model. On those simple tests, the two performed about the same. Asked what ER 2 can do that Flash cannot, Gopalakrishnan listed the benchmarks her team tracks: instrument reading, inspection and "a lot of pointing." Google's ER 2 announcement describes an instrument-reading evaluation covering ten categories, including digital displays, rulers and liquid thermometers. She described the relationship with the main Gemini models as "a collaboration," not a competition, with "a lot of data upstreaming." Even as general models get smarter, she said, specialized versions can still be fine-tuned to do slightly better at particular robotics jobs.
Why release the reasoning model first? Gopalakrishnan said ER is "much closer to Gemini itself" and therefore more mature. Actions are "farther out into the frontier," with research still needed on how to make them useful and serve them widely. The scale of use surprised her. Testing with trusted partners had not prepared the team for what happened once the model was on AI Studio. As a researcher, she was also struck by how academics picked it up, including in benchmarks that paired it with VLAs or even tried to use it for low-level control. Wide availability makes the model easy for outsiders to benchmark, which she called "free information for us" about its strengths and weaknesses. Asked to point to a project she liked, she mentioned a Boston Dynamics Spot robot reading instruments and handing out snacks. Google's launch materials show a Spot fetching popcorn after a spoken request, with ER 2 calling the robot's existing navigation and manipulation software rather than replacing its control system.
Six seconds to decide
The demo discussed in the episode also recorded timing. From a prompt like "move the banana to the plate" to the model's tool call took around six seconds, which seemed slow for the real world. The question put to Gopalakrishnan was whether this was the long-term design or just a convenient starting point.
Gopalakrishnan said it depends on the job. On an assembly line, "you can't loiter." At home, if she asks a robot to fold laundry, "I don't really care about how fast it is folding laundry. I care about how well it is doing the task. I can wait." She added that scale matters in robotics as it does in language models. The most advanced abilities are likely to appear first in the largest models, and larger models are naturally slower. She expects robotics to copy the Flash-and-Pro split of Google's chatbot models, plus an on-device tier, because even Flash runs in the cloud and needs a good network connection. Robots may be deployed far from good connectivity, on "some remote robot in like Michigan."
The host suggested this could reverse the usual expectation: robots too slow for factories might still be fast enough to do household chores overnight. Gopalakrishnan held to her earlier view that commercial settings will come first. They are more controllable places to try out new intelligence, while homes are more complex and face a higher safety bar, "given toddlers and stuff." She also rejected the idea that slow models will simply be sped up later. "I think we are going to simultaneously make the fast and the slow models and then distill them to one," she said. Distillation means training a smaller model to copy a larger one's abilities. She added that some abilities may appear only in large models and will have to be brought down to smaller ones afterward.
About three minutes of memory
Google's API specifications list a 131,072-token input limit for ER 2, the "128K" figure discussed on the show. Tokens are the small chunks of text, image or sound a model reads. Asked what that means for a robot, Gopalakrishnan estimated "about three minutes of memory." She noted that the figure depends on how video is broken into tokens and what other information shares the space. Google's streaming documentation, for example, accepts images at up to one frame per second. She said a long recording of a robot doing many things will not fit into ER today. For offline analysis of long recordings, she expects the Pro models, with their longer context, to be more useful.
Her advice on managing that budget drew on how people remember. If she is cooking, she can recall the moment she flipped an egg in detail. But much of the time, a running narrative of what she did, "even in text," is enough. Text is more compressible than images, so robots may keep recent or important moments as images and summarize the rest in words. What to keep, she said, depends on what is relevant to the task. She linked this to the hand-off between ER and the VLA: ER parses a human's intentions and passes them to the VLA through modalities people can read and understand.
How the planner talks to the body
Any tool a developer describes can be given to ER 2, as the demo showed. Gopalakrishnan said ER is built on Gemini's general tool-use abilities. She guessed, while saying the team would need to test it, that it handles widely known robot interfaces better than unusual ones, because there may be no standard for robotics APIs. Some researchers are trying to have large models drive a robot's joints directly. For now, she said, robot-specific foundation models still do that job better, though the larger models show "impressive capabilities in orchestrating these robots."
The host asked how ER tells the VLA which yellow object to pick up when a banana and a lemon are both on the table. Gopalakrishnan answered that her reply was "only true at this point." The research area is called steerability: the ways a VLA can be guided. Today VLAs can be steered with language and by pointing to objects. That is not the same as giving them pixel coordinates, she said: "I'm literally just doing pointing." Researchers are also testing whether VLAs can be prompted with video demonstrations. ER sometimes gives body-level directions too, such as "turn around" or "look to your left," and the VLA responds, even though it was not specifically trained for much of that. With foundation models, she said, "you can be quite surprised at what they do."
She also named a limit. "If the modality between the ER and the VLA is just text, a lot of things can be lost in there," she said, so the team wants that interface to grow "as rich as possible over time."
One model from fingertips to feet
The host wondered whether a model built on a language-model base could run fast enough for the split-second balancing a humanoid needs, or whether the robots handle that themselves. Gopalakrishnan pointed to the humanoid races in China. Controllers can already move a body, but "the very low level controllers often don't have a semantic brain," so they cannot work out whether the robot is doing what a person asked. She argued that this cannot be neatly divided into one system for meaning and another for movement. Only by reasoning about the whole body together with sensory information about the surroundings, she said, can a robot do many different tasks. Her conclusion: "I think that is the right design." She did not describe the lower layers as empty. The VLA controls the whole body, she said, but "there are more decisions being made about how to stabilize and how to achieve the targets of the VLA."
Google's announcement of Gemini Robotics 2 shows the same model checkpoint controlling three different embodiments: the Apptronik Apollo 2 with SharpaWave hands, the Apollo 2 with Inspire hands, and the Franka Duo with a Robotiq gripper. Its charts report success rates on a wide variety of whole-body and dexterous manipulation tasks. Each bar is an average over multiple tasks within a skill category, while for multifinger tasks the page shows individual task performance.
Where it breaks
The host asked whether most failures come from the action model rather than the reasoning model. Gopalakrishnan gave several sources. The VLA is being developed with partners, which she named as Apptronik, Agile Robots and Boston Dynamics. Agile Robots announced its partnership in March 2026. Robotics still has an "embodiment problem," she said, because VLAs learn from data collected on particular robots. Taking a new robot body and having it perform many tasks reliably with no extra training has no strong precedent, she said.
She identified the hand-off between the two models as another weak point. One model performs the task while the other watches to decide when it is finished and what comes next, which adds delay. "A lot of failure cases often also come from how to do the orchestration really well," she said. Then there is compounding: with ten tasks in a row, an ER model with one success rate and a VLA with another, "you are basically compounding the error." As an illustration with invented numbers: if each model got each step right 95% of the time, a ten-step job needing both models at every step would finish without a mistake only about 36% of the time.
Big-picture and fine-grained control
Near the end, the host asked how she would want to deal with many robots at once: as separate characters, or through one assistant that manages them as a fleet. Gopalakrishnan said the way people interact with robots needs to be "highly multimodal," and that users need "both high level and low level access to robots." She called this "probably the only way that we can build towards a very safe and flexible future." Some jobs will involve several robots handing work to each other. She pointed to the Gemini Robotics 2 demonstration in which a two-arm robot and the Apollo humanoid split a task, one stopping and the other taking over. Other jobs involve a single robot, where she might just talk to it or might need to give much more detailed feedback. For safety and interpretability reasons, she said, low-level access may need to remain.
The discussion compared this to software apps where an AI assistant sits beside the regular controls. Most of the time it is easier to just tell the assistant what to do, but it is frustrating when the app won't let a user adjust the details, because the AI sometimes does not get it on its own. The same paradigm, the discussion suggested, could carry over to robotics. Gopalakrishnan's example was her desk. "I cannot tell a robot to clean my desk because I like my desk done in a certain specific way," she said, so she wants to be able to give it more feedback and personalize it to her liking.