Asked whether one kind of data will end up training robots, Keerthana Gopalakrishnan, research lead for Gemini Robotics at Google DeepMind, declined to pick one. "I think it might also very much look like it's a mixture over time," she said on The Cognitive Revolution, in an episode published October 3, 2026. She compared the main data sources by how far each can scale and how precise it is. She gave a similarly open answer on a central debate in the field: whether robots need world models.
The case for one dominant source
The question came from a talk by Jim Fan of NVIDIA, which was credited in the discussion with inspiring a couple of big questions. (The episode's text renders the name as "Jim Phan"; the NVIDIA researcher is most likely Fan.) As the discussion summarized it, Fan expects first-person, or egocentric, video of people doing things to become the main training data for robots. Teleoperation, where a person remotely steers a robot and the robot's exact movements are recorded, would shrink to a tiny share. Wearable robot hands and similar devices would also give way to first-person video. After that, in this telling, simulation takes over: Fan expects the remaining problems with simulated training to be solved, making simulation the route to enough data.
The same talk, as described, argued that robot models need to think ahead. A system that looks at the present, reasons about it, acts and repeats the loop is not enough. A robot needs some forward-looking simulation of how things will unfold, so it can measure the gap between what it expected and what actually happens. That idea is closely related to world models: learned systems that predict how the world will change. The discussion also raised a counterpoint: large language models have come a long way predicting one token, or word fragment, at a time without explicit world models, and robotics might do the same.
"The jury is still out"
Gopalakrishnan would not take a side. "I think the jury is still out there," she said. Roboticists, she argued, are "not emotionally attached to one method or the other method. You're emotionally attached to the problem itself," and will pick whatever solves it. She called world modeling "a very interesting area for robotics," both for policies (the part of a system that chooses a robot's next action) and for simulation. She also pointed to vision-language-action models, which take in camera images and instructions and output robot movements, and to the agentic approaches now emerging in robotics.
For her, the open question is part of the appeal. "It is still very early that the recipes haven't stabilized," she said. People hold very different ideas about how to solve robotics, "and all of them could be right." Until the problem is solved, she said, it is hard to say definitively which approach is the answer.
One example of the world-model approach, not discussed in the episode, is a research system called DreamZero. Its authors built a "world action model" that predicts future video and robot actions together, swapping in real camera frames as the robot runs so later predictions respond to what actually happened. In tests on AgiBot robots, it averaged 62.2% task progress on tasks present in its pretraining data, evaluated in zero-shot environments with unseen objects, against 27.4% for the strongest pretrained vision-language-action model it was compared with. Those figures measure partial progress, not finished tasks. The authors also list remaining problems with sub-centimeter precision, short visual memory and computing cost.
Scale versus precision
On data, Gopalakrishnan said predictions should be tested rather than declared. "I think everything needs to be backed by results," she said. She did not dismiss Fan's views: "it's great that he has very strong views on certain types of data." But she pointed to language models, which are trained on all kinds of data, each serving a different purpose.
She suggested imagining each data source on a chart, with scale on one axis and precision on the other.
Teleoperation sits at the precise end. It is "exactly the robot's data," she said, and useful for grounding a model in the robot's own controls. But it is not very scalable, and it is not "future proof": as a robot evolves, old teleoperation data becomes less useful, even with good cross-embodiment, meaning the ability to carry skills from one robot body to another.
UMI-style data sits in the middle. The name comes from the Universal Manipulation Interface, a research system in which people perform tasks with handheld grippers instead of a robot. Gopalakrishnan said this approach allows "much cheaper scaling" because no robot is in the loop. And because sensors on the device record the motions, the action labels (the record of what the hand actually did) come from those sensors rather than from vision, making them "a bit more accurate." The tradeoff, she said, is that the sensors make it less scalable than plain human video.
The original UMI paper shows both sides. Its handheld grippers carry wrist-mounted cameras, and their motion is reconstructed by combining camera images with inertial measurements. For one cup-handling task, three people collected 1,400 demonstrations across 30 locations in 12 person-hours. Robots trained on that data succeeded in 43 of 60 trials in two environments they had not seen. The authors also note that the bulky grippers, with few degrees of freedom, limit what people can demonstrate compared with bare hands.
Egocentric human video is the most scalable source, and Gopalakrishnan said it is useful for "getting broad semantic understanding" of what to do. Its weakness is precision. If a robot needs to control very fine hands, she said, video gives no direct signal of where a person's hands are; that has to be estimated from images, "and vision has a lot of errors." The movements then have to be put onto a robot, which adds error. And people differ: "Some humans are small, some are large," she said, so the way they carry out the same task varies.
Hardware moves the balance
Gopalakrishnan's conclusion was that robot training will probably be "closer to a mixture," and that the right mix will not stay fixed. New hardware changes each approach differently. As "very good UMI hardware comes on the market," she said, the amount of that data that can be collected goes up. "So I don't think it's wise to take a very principled view about this because all of this will change as new evidence emerges and better hardware emerges," she said. "And you probably need all of all types of data."