[BidClub_]
The Cognitive Revolution · · 92 min

One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids

Nathan LabenzKeerthana Gopalakrishnan

RoboticsAI & SoftwareTechnical
YouTube ↗
TL;DR
  • Keerthana reads China’s viral Robot Olympics as genuine locomotion progress but a poor proxy for near-term robot utility. Running on flat, rigid terrain is unusually simulation-friendly, while useful manipulation breaks down around contact, friction, deformable objects, and unpredictable environments. Her blunt calibration: “Many of my friends are very productive, but we don’t run faster than Usain Bolt.”

  • Robotics remains at its “GPT-2 moment” because impressive demonstrations have not yet become portable, reliable intelligence. GPT-3-level progress would require learning from multiple examples across many tasks and a brain that transfers across humanoids, arms, and other bodies; today, competence still depends heavily on the exact robot and setup. One-shot imitation is promising, but copying a demonstration is not the same as generalizing when the scene changes.

  • Gemini Robotics 2 separates embodied reasoning from physical execution across three models. Gemini Robotics ER2 is the higher-level, tool-using “brain”; Gemini Robotics 2 is the vision-language-action model controlling motion; and Gemini Robotics On-Device 2 is a smaller local version. The architecture lets a large model interpret intent while specialized action models handle bodies, but additional model handoffs create latency and failure risks.

  • Controlled commercial environments remain the likeliest early deployment wedge, even though model latency may matter less at home. Nathan emphasizes repeatable factory and warehouse tasks and professional servicing; Keerthana agrees that commercial use cases are more controlled than homes. A six-second pause is untenable on a conveyor belt but irrelevant if a robot folds laundry overnight. Keerthana expects fast and slow models to develop together, with capabilities transferred from larger systems into smaller ones.

  • Cross-embodiment is improving rapidly, but “one brain, any body” is still an aspiration rather than a solved capability. Google DeepMind’s models now control humanoids “from fingertips to feet,” and its on-device program reportedly adapts tasks to new bodies with roughly 200 examples. Yet Keerthana sees little precedent for taking a new embodiment fully zero-shot to many tasks at very high reliability.

  • Hardware has moved faster than Keerthana expected, shifting the remaining work toward reliability, repeatability, cost, and safe force control, while dexterous hands approach a capability bottleneck. In roughly a year and a quarter, the frontier advanced from grippers to multifinger control, garbage-bag tying, and whole-body manipulation. The Shadow hand, she says, may lift about 20 kg. The remaining commercial work is making hands repeatable, durable, affordable, and safe around delicate objects and people.

  • Keerthana refuses to declare either world models or any single data source the winning robotics recipe. Teleoperation is accurate but expensive and difficult to scale; instrumented UMI demonstrations scale better while retaining precise sensor-derived action labels; egocentric human video is abundant but noisy. Her recurring principle is empirical: forecasts should be tested against results rather than treated as settled, because robotics remains early enough that several architectures and data mixtures may prove useful.

Digest · the substance, structured for research

1. Viral robot races showcase the easiest part of embodiment

  • Keerthana first encountered China’s Robot Olympics through her father and friends in India, who were struck by humanoids apparently running faster than the fastest people. The demonstrations captured the public imagination because running is an intuitive benchmark—and because robots have progressed dramatically since DARPA-era machines struggled to walk through doors without falling.

  • Nathan’s pushback—worth keeping—is whether speed is useful rather than merely spectacular. Keerthana allows that fast locomotion might matter in defense, but her team is focused on robots helping people in the physical world: “Many of my friends are very productive, but we don’t run faster than Usain Bolt.” Leg speed is rarely the limiting factor in valuable work.

  • Locomotion has advanced faster than manipulation because flat ground and rigid bodies can be modeled relatively well. Controllers can learn motion in simulation, often through imitation learning followed by reinforcement-learning fine-tuning, and transfer it to a physical robot; contact-rich tasks introduce harder physics. Picking up a wooden cube is tractable, while folding cloth exposes friction, deformation, and interactions that modern simulators still represent poorly.

2. The research frontier has moved from grasping to whole-body intelligence

  • Keerthana remembers when basic pick-and-place was a serious research problem. Foundation models made many of those behaviors comparatively easy, pushing researchers toward humanoids, multifinger hands, and coordinated whole-body manipulation. Her team began working on humanoids around two years earlier, when even grasping an apple was difficult.

  • Gemini Robotics 2’s meaningful advance is not simply walking or moving quickly; it can squat, reposition itself, and manipulate unfamiliar objects in a closed loop. Keerthana points to a video in which a humanoid changes posture while trying to retrieve a watering can—behavior she would have found surprising just a year earlier.

  • “Whole-body control” means the vision-language-action model considers the humanoid from fingertips to feet while processing the surrounding scene. Lower-level mechanisms may still contribute stabilization, but the semantic controller is not merely issuing commands to an isolated arm: it coordinates bodily actions in service of a user’s task.

3. Reliability and generalization remain separate axes

  • Keerthana frames skill and generalization as “orthogonal axes.” A narrowly engineered system can already perform one or two tasks extremely well, but the cost of executing many narrow tasks can become roughly the same as the cost of executing the first or second. A general foundation raises performance across many tasks and makes subsequent specialization cheaper, echoing the transition from task-specific language models to GPT- and Gemini-style foundations.

  • Nathan presses for a METR-like curve connecting task complexity to 50%, 80%, or 99% success. Keerthana does not offer a clean curve, but says pick-and-place is approaching a strong, usable signal. Her larger point is that robotics cannot optimize only for task success: it must also preserve the broad understanding that supports adaptation.

  • Physical deployment sharply raises the reliability bar. A language model can omit something and be corrected; a robot can break an object, create a hazard, or damage itself. Tasks with cheap retries are therefore much easier: a robot can undo a LEGO mistake, but after dropping an egg it has created a slippery mess and a second problem to solve.

4. One-shot demonstrations are prompts, not proof of generality

  • Recent systems can learn from a handful of demonstrations—or apparently one—but Keerthana places these results on a continuum with ordinary prompting. A foundation model may be guided by language, an image with an indicated object, or a video demonstration. In-context learning reduces deployment time because it avoids task-specific retraining, but it does not abolish the underlying generalization problem.

  • The critical test is what happens after the exact demonstration changes. These models can memorize and reproduce what they were shown; the harder question is whether they adapt to a different scene, object, or arrangement. Pick-and-place benefits from abundant prior data, while “show it once, then tie a garbage bag” remains a much stronger challenge.

  • Keerthana therefore still calls robotics “GPT-2.” Reaching a GPT-3 analogue would require learning from multiple examples across many tasks and substantially stronger transfer across bodies. A brain that works only on one robot in one configuration is not yet a general brain in the way software behaves consistently across phones, computers, and operating systems.

5. Gemini Robotics 2 divides reasoning, action, and local execution

  • Keerthana describes Gemini Robotics ER2 as a tool-using embodied-reasoning system built on the Gemini Flash models, oriented toward understanding images, video, semantics, and robotic situations. Gemini Robotics 2 is the VLA that converts user intent or ER2 guidance into action, while Gemini Robotics On-Device 2 is a smaller model designed to run locally on the robot.

  • Nathan says his coding agent connected ER2 to a browser-based 2D scene with simple grasp-and-place tools. The model understood requests such as moving a banana to a plate. He also tested an object that moved before grasping; the model failed, and then received another image with the object in a different place. In his small test, Flash and ER2 performed mostly the same overall, and both did well.

  • ER2 was released via API first because it is closer to the mature Gemini stack. Action models remain farther from broad deployment and require more research into reliability and serving. Opening ER2 beyond verified partners also produced a much larger stream of feedback, including academic benchmarks and experiments using it for unexpectedly low-level control.

6. Latency and memory force a hierarchy of robotic brains

  • Nathan observed about six seconds between asking ER2 to move an object and receiving a tool call. Keerthana’s answer is workload-specific: a conveyor cannot wait, but a homeowner may not care how slowly laundry is folded overnight. Larger models may unlock advanced capabilities while smaller or on-device models supply responsiveness, local execution, and operation without dependable cloud connectivity.

  • She expects fast and slow systems to coexist rather than arrive sequentially. Some abilities may emerge first in large models and then be transferred into smaller ones. Cloud deployments can run on distant servers—for example, in Michigan—while on-device models address applications that need local execution.

  • Nathan describes ER2’s 128,000-token context as roughly three minutes of densely represented robotic history, depending on frame sampling and tokenization. Keerthana’s context-engineering principle is “very dense memory”: retain raw visual detail when it matters, but summarize long activity into compact text when the exact frame is less important. A cooking robot may need to remember that it flipped an egg, not preserve every pixel of the motion.

7. Rich interfaces matter more than a universal robot API today

  • ER2 can be given tools much like a digital Gemini agent, but robotics lacks a standard interface equivalent to the Model Context Protocol. Every body exposes different controls and APIs, so unfamiliar embodiments may require additional engineering. Keerthana says current robot-specific models remain better at direct control for now, with larger reasoning models orchestrating them.

  • The interface between ER2 and a VLA is already broader than raw coordinates. A VLA can respond to language, pointing, images, or video demonstrations; ER2 can instruct it to look left, lower its head, return, or interact with a referenced object. Keerthana’s direction of travel is a richer connection in which text is only one communication modality.

  • Nathan’s banana-versus-lemon example exposes the design question: should ER2 name the object, point to it, describe its shape, or provide pixel coordinates? Keerthana does not declare one permanent answer because “controllability” is an active research area. The appropriate abstraction will evolve with what action models can reliably understand.

8. Cross-embodiment gains are real, but orchestration compounds errors

  • Google DeepMind is developing VLAs with Apptronik, Agile Robots, and Boston Dynamics. Keerthana says the field has made “a big jump” in cross-embodiment transfer, yet fully zero-shot deployment on a new body across many tasks at high reliability still lacks a meaningful precedent.

  • The on-device program demonstrates the value of a common physical foundation: with roughly 200 examples, testers can teach multiple tasks on different bodies because the model already understands much of the physical world and has prior control experience. Most embodiments nevertheless cluster around familiar designs—hands similar to human hands, vectors in space, bimanual hands, and humanoids—while hands mounted on drones sit in the exotic long tail.

  • Nathan points out that sequencing ten subtasks would create compounding opportunities for error. Keerthana emphasizes that completion detection, transitions, and deciding when to replan can fail even when reasoning and motor execution work well in isolation.

9. Hands are improving faster than their economics

  • In roughly a year and a quarter, Keerthana updated her own hardware forecast. Tasks once near the frontier for grippers can increasingly be performed by multifinger hands, while those hands add abilities grippers never had—garbage-bag tying, precise finger motion, and manipulation integrated with posture.

  • The market now offers several capable hands, but reliability, repeatability, and price remain unresolved. The Wuji hand, Keerthana says, is comparable to a ten-year-old child’s: human-sized but somewhat weak. The Shadow hand may lift about 20 kg and has been seen opening jars, though torque and success on tightly sealed lids remain uncertain.

  • Soft robotics matters where compliance and tactile feedback improve manipulation, but Keerthana again centers practical criteria: “What is repeatable, and what is durable?” Instrumented gloves and UMI-style systems are interesting because they can capture demonstrations and sensor-derived action information; force sensing also lets models join delicate parts without crushing them.

10. Safety is a capability spanning mechanics, perception, and intent

  • Keerthana rejects a simple capability-versus-safety tradeoff: “People will not use dangerous robots and dangerous agents.” The most useful systems must follow instructions without causing chaos or violating constraints. Nathan and Keerthana both expect commercial sites to precede homes; Keerthana emphasizes that commercial environments are more controlled, while homes contain children and carry a higher safety bar.

  • Robotics adds operational failures distinct from malicious behavior. A humanoid can injure someone simply by falling, losing a sensor, or misjudging contact. In one test, someone put a basket over the robot’s head; the right response was not to continue blindly, but to recognize the obstruction and request its removal.

  • Safety mechanisms range from hard emergency stops that cut power—potentially making an unstable robot collapse—to soft stops that freeze it while maintaining balance. Better force sensing and compliant control add another layer. For Keerthana, safety is a “full-system thing,” extending from mechanical design and emergency controls to perception, planning, and high-level adherence to human intent.

11. Humanoid interaction, world models, and data recipes remain open

  • Gemini Robotics 2 introduced more natural human-robot interaction: gestures are generated in context rather than individually scripted. During filming, a robot could discuss its own task, acknowledge where it might have erred, and gesture while conversing. Keerthana calls this “semi-emergent”—deliberately developed as an interaction capability, but not programmed gesture by gesture.

  • Human likeness raises both attachment and expectations. New observers can forget that a humanoid is a machine, while judging its mistakes more harshly than those of an obviously mechanical arm. Keerthana argues humanoids are scientifically valuable because they expose three frontiers simultaneously: whole-body control, multifinger dexterity, and natural human-robot interaction.

  • On predictive world models, her honest answer is that “the jury is still out.” Robotics researchers should be emotionally attached to the problem, not a favorite method; VLAs, action models, and explicit forecasting systems may all contribute. The field’s “recipes are not stabilized,” which makes empirical comparison more valuable than architectural certainty.

  • The same agnosticism governs data. Teleoperation is precise but costly and difficult to scale, and may become less useful as robots develop; sensor-rich UMI demonstrations are more scalable while retaining action labels; egocentric human video is abundant but noisy about exact hand and end-effector motion. Keerthana expects a mixture that changes as hardware improves, with users ultimately needing both high-level delegation and low-level control when personalization or safety demands it.

Full transcript

1. Humanoid Olympics and simulation

Nathan Labenz

Today, I'm glad to welcome Keerthana Gopalakrishnan, a research scientist at Google DeepMind and the research lead for Gemini Robotics, returning to the show for her 4th annual appearance.

One of the biggest questions in artificial intelligence today is how soon AGI will become broadly useful. Artificial intelligence is already significantly affecting the world, even in its purely digital form. But many of the most aggressive predictions about the future of AI suggest that robotics will soon reach key inflection points.

2. Sponsors: Deepgram Flux TTS | OutSystems | Claude

For example, the report “AI 2027” predicts that humanoid robots will become useful sometime in mid-2027. By 2028, it suggests, we may have a self-sustaining robotic ecosystem in which robots build more robot factories, which in turn produce more robots, leading to unprecedented exponential economic growth—but also creating a serious risk of an AI takeover.

So, how does this vision accord with the reality of robotics research today? We begin with a discussion of the extremely viral Robot Olympics, which will take place this summer in China. It was very interesting to hear Keerthana’s opinion about what it means that humanoid robots can now run faster than the fastest humans.

Her perspective, as you’ll hear, turns the usual U.S.-China dichotomy in AI on its head. Although the videos are obviously impressive, she says that skills such as running on a flat track are actually relatively easy to teach in simulation. More importantly, running speed is not really a limiting factor for the usefulness that robots can provide in the workplace.

Her team at Google is therefore more focused on practical value. From there, we move on to Gemini Robotics 2, a set of 3 models that Google released this summer. Gemini Robotics-ER 2, with “ER” standing for embodied reasoning, is a system for managing robotics based on Gemini Flash and available through an API that allows developers to define capabilities and make them available to models as tools, just as we do with digital agents.

I experimented with it and found it extremely capable. The other 2 models, Gemini Robotics 2 and Gemini Robotics On-Device 2, convert these higher-level tool calls into robot actions. They can now control everything on a robot, from its fingertips to its toes, across a wide range of form factors, although they are currently available only to trusted testers.

Having understood how these models work, we zoom out again and ask Keerthana how she understands the progress of robotics in general. On the one hand, like many other AI researchers lately, she says she has been surprised by the pace of progress. At the same time, she still feels that robotics remains in its GPT-2 era.

We’ve seen a number of recent robot demonstrations in which robots learn new tasks with the help of several, or even just 1, human demonstration. But Keerthana says that the range of tasks that can be taught this way, and the generalization of robotic models across different embodiments, still aren’t strong enough to ensure the versatility and reliability required for real-world use cases.

Near the end, we discuss recent progress in hardware, where work will probably unfold first at large scale, regardless of whether affordability, safety, consistency, and reliability ultimately become limiting factors for the consumer market. We also consider the future of the industry—in particular, whether we can continue to run robotics models derived from an LLM in a loop, as most modern systems do, or whether we’ll need some kind of world model, as humans use to create systems that can anticipate and interact smoothly with the world.

Regarding these questions and the time frames for key milestones in the field’s usefulness, Keerthana is not someone who speculates wildly. For her, these are empirical questions that have to be resolved through research and experiments. But after this conversation, it seems clear to me that robotics will either soon reach key inflection points, or some of the most aggressive timelines for AI transformation will have to be pushed back.

It’s nice to have you here. Great to see you again. It’s been about a year. Many things have happened. I wanted to start with something interesting. Recently, the Robot Olympics were held in China, and the videos obviously spread across the entire internet. People laughed, and sometimes, I think, even sympathized with the robots. I’d like to start with your observations and conclusions, given the depth of your knowledge on this topic. What surprised you while watching the videos from the Robot Olympics?

Keerthana Gopalakrishnan

I didn’t see them until my dad started talking about it and said, “Wow, did you see these running humanoids?” Then there were all these videos about humanoids running and falling. I also saw friends in India who had participated, and they shared various student posts. They said, “Wow, look at this. They’re building robots in China that can run faster than the fastest people.”

So, I think that, in a way, this competition, although it’s not very strange from a research point of view, really captured the public imagination in a very expressive way. I also think that this is a benchmark, right? Running is a benchmark, and people try to see where the robots are and where the gaps are.

If you look at videos of people who made these collages from the DARPA Robotics Challenge about 20 years ago, the robots could barely walk, open a door, and then fall over. And here we are. It’s amazing progress, and it’s also progress that depends very strongly on the public imagination.

The fact that a robot can run faster than Usain Bolt is, on the one hand, crazy progress. It’s insanely impressive.

Nathan Labenz

But one question that I have to press you on with this subject is: Is this valuable? Are the Robot Olympics simply—do you know—their main goal to capture the public imagination and show something cool?

Do we want robots that can run at 20 miles per hour in humanoid form? Do you think about this in your work and say, “I would want my robot to run faster. I need to work on this”? Is this just an extra topic?

Keerthana Gopalakrishnan

I think that, first, there are people who are working on research, and secondly, there are people who are trying to speed up this research. For better or worse, locomotion is research. There is locomotion, and there is manipulation.

Locomotion has advanced significantly, especially in this area, where modeling can go a long way. Many of these models are trained in simulation, and because of this, locomotion has progressed significantly. So, this is a demonstration of the current state of humanoid movement, such as bipedal movement.

I think there are many ways to look at this. I’m sure that, maybe for some defense applications, this is useful, but I think we’re very focused on making sure that the work helps people and enables useful things in the physical world.

I think I’m a very productive person. Many of my friends also work very productively, but we don’t run faster than Usain Bolt. I think that this is useful, but—

Nathan Labenz

That’s it. The bar is so huge, isn’t it? Different people are trying to advance in very different directions.

3. Outro

Tell me more about how to run into a wall during a sprint and break into pieces. It’s clear that they don’t train them like that constantly, because their physical hardware would wear out very quickly if they conducted all their training that way. So, based on this, I thought: Obviously, they probably do a lot of training in simulation.

Tell me more about why such tasks are trainable in simulation, and which tasks are still difficult to teach in simulation.

Keerthana Gopalakrishnan

Yes. Sim-to-real is what we should pay attention to. In robotics, sim-to-real means the gap between whether I can teach something in simulation and then do it in reality.

As we know, especially from agents and other systems, the things we can learn in simulation are probably the first things to be solved in this sense. In locomotion, you can often model the ground or the wall as something that comes into contact with the robot, and these are often very flat surfaces. This makes the task very suitable for simulation, and then you can do a variety of things.

But with manipulation, you’re really interested in contact. When does contact happen, and how do objects behave? This is where modern simulations begin to break down. For example, folding fabric is more difficult because there is friction in the fabric.

Even in simulation, many problems involving picking up and placing objects can still be solved, because the contact dynamics are often predictable to some extent. The stiffer the body—whether it’s the floor or, for example, a wooden cube—the less deformation there will be, so it will be easier to simulate.

Nathan Labenz

Yes. Contacts are essentially physics that is more difficult to model in simulation. So, where the physics is easier to simulate, the simulation can encompass it.

Regarding the artificial intelligence component of these robots, what conclusions would you draw from what we saw? Of course, I don’t know how they did it, but it seems that there’s no obvious ingredient of reasoning about what’s happening.

I don’t have a clear idea of which architectures were used and which were not, or whether all the models used to control the robots were on the device. Maybe some of them were running off the device. How do you imagine they usually do this? How high in the stack, in your opinion, did they work?

Of course, there are many systems working on stabilization to make sure that the robot doesn’t fall over at any moment.

But is this all? How would you do it?

Keerthana Gopalakrishnan

I appreciate the question, because we're going to move to Gemini Robotics, where there is also a much higher level of this architecture: scene reasoning and understanding. I suppose that all of this was somehow missing for the purposes of these Need for Speed demos. I didn't work on those demonstrations, so I can only guess.

4. Generalization versus task mastery

There is a very wide published corpus of work on locomotion and simulation, and I think these are very good RL controllers, probably trained through imitation learning and then fine-tuned with the help of good RL, which is suitable for whole-body control and for that particular task, where they're aiming at very high speeds. They may also have some navigation, although I think running a race is quite easy.

Nathan Labenz

You mentioned that the goal is to work on useful things for people in the world. Obviously, speedy leg movements usually aren't a limiting factor for what humans can do. In the modern world, we live and work next to robots. Is there something you would delegate to a robot at a practical level, where you think it's skilled enough that it takes that task off your plate and lets you focus on research interests and all these other things? In other words, are there tasks where robots win?

Keerthana Gopalakrishnan

Yes. The strange thing is that I'm actually a researcher, right? So if robots win at some tasks, I'm going to work on tasks that robots cannot do. Movement is the limit.

I remember when I started working on manipulation, even choosing and grasping an object was very difficult. We didn't have the right algorithms, but now, with base models, enough of these things are easy. Therefore, the research boundaries have shifted toward newer things for humanoids.

Nathan Labenz

So if you look at state-of-the-art robotics, right?

Keerthana Gopalakrishnan

Yes. It's about whole-body control and whole-body manipulation. This is not something that someone has demonstrated before, especially whole-body manipulation with generalization across different object types. It's not just very fast running or the ability to take steps; it's the ability to squat and manipulate objects, doing all of that in a closed loop, with generalization across different object types.

I would have been a little surprised if, a year ago, you had told me that we would be doing this now. We've been working on humanoids for about 2 years. To begin with, it was just an attempt to grab an apple, and even that was very difficult. Now you can have it grasp objects. I posted a video on Twitter where I play with it using a watering can, and it sits and squats in different places and tries to lift it. Things are going well enough so far.

Nathan Labenz

Maybe another way to formulate the question is from the point of view of the METR curve. I'm sure you're familiar with the famous curve for the duration of tasks, which is usually depicted as showing how, over time, models can perform tasks of increasing duration, measured by how much time a person needs to execute the task.

5. Whole body humanoid control

But there is also the factor that a robot—or artificial intelligence in general—may perform a task with 50% success, or it may have a version that performs it with 80% success. Obviously, one big difference between using an LLM and deploying robotics is that 50% success in the real world probably isn't enough, unless you have a genuinely fault-tolerant task. Whereas on my computer, I can say, “Something was missed? Well, okay.” But if the robot misses something, it breaks, and we have a problem.

If you tried to describe where we were recently and where we are now using a curve similar to METR's, which shows how complicated the tasks are that robots can do well enough to be useful, what things can they perform more than 99% of the time? At what level would they be approximately 80% today? At what level would they be approximately 50% today?

6. Gemini Robotics 2 architecture (Part 2) (Part 1)

Keerthana Gopalakrishnan

I think a lot of pick-and-place is close to being solved, with very good reliability, and close to being useful. However, I think there is a balance we need to achieve in robotics and AI in general. You can create a very narrow AI that does one thing very well, but then you lose the whole point; you're not working on generalization. You can consider generalization and skill as somewhat orthogonal axes.

Work on generalization is like lifting all boats, raising a wave for all the boats. If you imagine tasks laid out along this spectrum, having very broad common sense and a very general understanding significantly facilitates achieving skills in every task. This is also an important part of how we think about approaching this.

Even today, you can create a very narrow AI that is very good at one thing or another two things. But then the cost of executing N things becomes approximately the same as the cost of executing the first or second thing. Building very general basic models significantly facilitates rapidly acquiring skills.

I think we saw this even with LLMs, right? At first, people had specialized models that performed one thing, but now GPT and Gemini are very good at a wide range of tasks. Based on these general foundations, you can rise very high. I think the same thesis can be true for robotics.

Nathan Labenz

Yes, that's interesting. Last year, when I saw your story on the curve, one of the interesting moments, as far as I remember, was taking the base model you had at that time, conducting some precise fine-tuning for specific tasks, and showing the benefits of that fine-tuning.

Recently, I've seen several demonstrations that I definitely couldn't understand in practice, but online you can see videos from different robotics companies saying, “Hey, look, we have generalization from several examples,” or even, “We have generalization from one example.” A person performs the task once, demonstrates that it works, and then the robot can start performing it.

Of course, that's great if you can achieve it, but how, in your opinion, is this actually likely to work for deploying robots? What should I do? Are they going to make 100 demonstrations of a task? Do they even have to make 10? Can it be that simple? Or do you think we're at the limits of having many things work only on a one-shot basis? What investments should people be ready to make so that their humanoid is tuned to what most of them care about?

Keerthana Gopalakrishnan

Yes. I think there will certainly be a spectrum. The possibility of ICL, or in-context learning, where you show how to do a task and then the robot performs it, definitely reduces deployment time and execution time, right? You don't need any special fine-tuning, and you don't need to train it. You can do it in context, and this is definitely a very exciting development.

You can consider this like video prompting, or something similar. Even older models, for example, when I say, “Take the object,” and it's a novel object, what am I doing? I send language, and then I get very general behavior that I have never shown before. This is generalization.

Now, instead of prompting with text, you can prompt with an image. You can say, “Here's something,” and then circle what you want to manipulate. This is a hint for using the image. In a sense, it's also a hint for using video.

I think ICL results so far are on a spectrum of how you prompt a base model. You can prompt it through video, through language, or through an image. The question is what kind of generalization you can get. These models are also prone to memorization in a certain sense. If you show them exactly this, they will copy it. But the question is whether they can generalize or not.

Nathan Labenz

If you change the scene and other things, can they do it in an adapted way?

Keerthana Gopalakrishnan

Then there is a second question concerning the task itself: what difficulty level of task are you demonstrating? Pick-and-place models have a lot of data, and it's also easier to understand. But can you show a robot how to tie a garbage bag? Can the robot tie the garbage bag after you show it once?

I think it's still quite early, and we need to see how it develops.

Nathan Labenz

This may be a stupid question, but it could be useful for calibration. If you had to evaluate robotics today in terms of GPT-1 through GPT-6—or, as we discussed previously, if we may have reached the GPT-2 moment—where would you rate the field today?

I think GPT-3 is synonymous with in-context learning, but I'm going in a slightly different direction with robotics, where we have built-in instruction, perhaps in many cases even as good as in-context teaching. That exposes the flaw in my question, but if you had to draw an analogy with how many GPTs along that progression we have reached in robotics, where would you rate the field today?

Keerthana Gopalakrishnan

Still GPT-2, and here's why. For GPT-3, first of all, you need learning with multiple examples to make it work well for many different tasks. Secondly, robotics, unsurprisingly, is still very dependent on the environment, right?

If it only works on your robot with your specific configuration, is that really a general brain? Can I put this brain on your humanoid, on another robot, or on some other robot I'm working with, and just bring it over? If it's absolutely helpless in that situation or environment, is it really a general brain?

I think GPT would never have such problems, right? My phone or your phone, my computer, a Linux machine—it doesn't matter where you run it. Once you start it, it behaves the same way. You can expect very similar behavior.

But here, I think we are very dependent on what kind of robots you are influencing. Therefore, I think there is much more work needed to create a very general brain that can account for these factors.

7. Interacting with robot fleets

Nathan Labenz

Okay, maybe this is the perfect transition to Gemini Robotics 2. To begin with, just give us an architectural overview. For example, when you go to the webpage, there are 3 models released: Gemini Robotics-ER 2, Gemini Robotics 2, and Gemini Robotics On-Device 2. Describe how they relate to one another. Give me a general overview, and then I’ll get to more specific questions.

Keerthana Gopalakrishnan

You can imagine Gemini Robotics-ER as the brain of a system, a tool that can perform very general reasoning. It’s based on the Gemini Flash models, but perhaps more oriented toward robotics. It can perform many robotics tasks, and it’s also widely capable: you can ask it for general image and video understanding, as well as semantic understanding.

You can imagine Gemini Robotics as a VLA that controls the robot. It determines how the robot moves and what it must do, given a certain user intent. You can imagine the Gemini Robotics On-Device model as a much smaller version of what’s placed on the robot’s computer. It isn’t in the cloud, but it does something comparable. Essentially, it can do what Gemini Robotics does, but it’s a smaller model placed on the device.

Nathan Labenz

One question I have is about how this is done. This is based on Gemini 3.5 Flash, and there’s a lot connected to that, isn’t there? For example, do you have chain-of-thought? Do you have multimodal inputs? Everything is so different. Is this the long-term architecture on which you expect robotics to work?

One obvious problem is that language models, especially when they’re doing reasoning, take some time to produce an answer, right?

Keerthana Gopalakrishnan

This model is in the API. This is another topic worth discussing. Because it’s in the API, I was able to get a coding agent to use it and provide some tools, and I created small, 2D demonstrations where I could insert a small image of the scene. It was really just a small canvas in a browser.

Surprisingly, I could get pretty far with only a few prompts to understand what the model can do and how simple it is to use. I also recorded the timing. Sometimes, from a prompt like “Move the banana to the plate” to the tool call took approximately 6 seconds. That seems like a long time for the real world.

Nathan Labenz

I wonder if this is the basis of the approach you’re building now. Is this a paradigm that, in your opinion, we’ll ultimately see—perhaps just accelerated or modified somehow—to work adaptively in our environment?

Keerthana Gopalakrishnan

I think it all depends on the specific application. In certain applications, for example, if you’re on a conveyor belt, you can’t idle, and you need to do things very quickly. But let’s say I’m at home or somewhere else and I say, “Hey, robot, can you do my laundry?” It doesn’t really matter how quickly it folds the laundry. I’m concerned about how well it performs the task. I can wait.

There may also be a question of scale, which is a very important factor in LLMs and robotics. I think a lot of advanced capabilities will appear in very large models, and these models will naturally be slower than smaller models. So I think we’ll definitely need to find a compromise and a spectrum, and I think the Flash and Pro paradigms will also exist in robotics.

Even in robotics, you can consider this as a paradigm for devices. Even a Flash model can still work in the cloud, right? You need a very good network connection to communicate with it. The deployment of robotics can be very distributed. You can run it on some distant server, for example, in Michigan. So I think these paradigms will develop, but I think every scale will be relevant and useful for various types of applications.

Nathan Labenz

I wonder what you compare this with, because I usually think that work will first go to warehouses, factories, and commercial environments, where there is a lot of control and consistency in the tasks the robots need to perform. Of course, there are professionals who can service a large number of them.

All of these things historically forced me to think that robots wouldn’t come to my home until they had been deployed quite successfully in commercial settings. But what you just said pushed me a little to the other side. This could be insufficiently fast for a factory, but it could definitely be fast enough to take care of all your overnight tasks, and that could be the inversion.

Do you still think that, the last time we talked about it, you expected commercial use cases to come first? Is this still the case, and how does it relate to the question of reaction time?

Keerthana Gopalakrishnan

I definitely think so. Commercial use cases are probably more controlled, so they’re more suitable for trying various new intelligences compared with homes, which are more complicated and where you can’t control things. There’s also a higher standard that needs to be met for safety, taking toddlers and other factors into account. So I still believe in that.

As for speed, I don’t think it will be that we first create slower models and then make them faster. I think we’ll simultaneously create fast and slow models and then combine them into one. It’s quite possible that some capabilities will first appear in larger models, and then we’ll need to figure out how to transfer them to smaller models.

I just want to highlight the difference between the embodied model, Gemini Robotics-ER 2, and the base model, Gemini 3.5 Flash, on which it’s trained. One of the things that my coding agent did without my request was compare them. It gave them both something like, “Here’s the scene, and here are the tools available to you.” The tools were quite high-level in the small test I conducted—for example, “Grab an object and then place it.”

It even did one test inspired by famous videos, where an object moved just before I tried to grab it. The model failed, and then it received another image in which the object was in a different place. Those were all relatively simple things, and in the test I conducted, Flash and ER 2 performed mostly the same. In fact, both did well.

Nathan Labenz

How would you describe the task boundary for embodied reasoning that ER 2 can perform but Flash can’t?

Keerthana Gopalakrishnan

Many of the benchmarks we track are similar to reading instrument indicators. We think ER is much more capable, especially in robotics use cases. This looks like checking instrument indicators and following many instructions, but it’s also very similar to cooperation, right?

It’s not as if ER and Gemini are competing. There are many data streams, so it’s very much a joint effort to make Gemini models really good at robotics in general. If that isn’t sufficient, then the development path could look like this: even if Gemini models have more general intelligence, you can still create specialized models that are a little better for a specific robotics use case, if that use case requires more specific engineering or fine-tuning.

Nathan Labenz

I think the question consists in what you want and what you want to design. The model has a 128,000-token context window. How much is that in the context of robotics? I have enough intuition about how many pages of text that is, but I don't really understand very well how history is encoded episodically or how many images per second.

How should we think about this from the point of view of time? For example, how much does 128,000 tokens mean from the point of view of the maximum episode duration for a Gemini Robotics-ER 2-based robot?

Keerthana Gopalakrishnan

This is approximately 3 minutes of memory, and this depends on how you tokenize, how many tokens you use to represent things, and what other information you provide. However, I think it's about memory. You will see that Pro models have much larger context lengths, and they also become much better at fitting in larger amounts of information.

I think Gemini Robotics-ER 2 is not available now for a very long episode in which the robot does many things and you're trying to get something from that result. Since working over the length of the context is a constant problem, I think that, especially for use cases such as offline analysis, Pro models will be more useful.

Nathan Labenz

So 3 minutes would be if you had a very tightly packaged context? If you're giving advice to users of the API, I would aspire to more. I obviously can reduce the number of sampled videos, have fewer frames, and do all sorts of things, right?

I'm working on contextual engineering for robotics. Perhaps the question I really want to put to you is: What have you learned about context engineering for robotics that people should know? How should they think about which parts of the interaction history they need to save to have consistency? Maybe it's similar to LLMs, but I think there are some differences.

Keerthana Gopalakrishnan

Yes. I think very dense memory is important. There is something—I don't know, I think this is Daniel Kahneman, or someone from Thinking, Fast and Slow—but even for aspects such as memory, especially for something like, “What did I do in the past?”, is it worth keeping it in a very dense modality, like a picture, or can I use a lot of text summarizing what I was doing during that time?

For example, if I cook, I don't have confidence that, if I want to, I can go back to that time when I, I don't know, turned the egg or something like that. But often, information or a story about how I do things, and a summary even in text format, is quite useful or informative enough about what I was doing.

I think the context depends on what you want to save and what information is relevant to the task. Obviously, if you can remember everything, then that's perfect. But in practice, you need to be more effective in managing your RAM, right? You want to make sure that the most important information is presented. Sometimes this may take the form of information represented as text that is a little more concise.

This is also connected with the distinction between Gemini Robotics-ER 1.5 and VLA. ER is trying to discern a person's intentions and communicate them through text and other modalities to VLA.

8. Gemini Robotics 2 architecture (Part 2) (Part 2)

Nathan Labenz

Before we delve into VLA, why did you and DeepMind as a whole decide that this was the first robotics model released via API? And what did people create with its help?

Keerthana Gopalakrishnan

Gemini Robotics-ER 2 is much closer to Gemini itself in terms of capabilities, so it is more mature. Action models, in my opinion, are still much less mature, so we need to do a lot of research on how to make them very useful and serve them more broadly.

It was quite strange to see how people were using it. I saw that it had such widespread popularity, and the team did really beautiful work there. One thing we found out is that, although we spent a lot of time testing with verified partners and similar projects, as soon as we made it more widely available in AI Studio, usage was on a very different scale from anything we learned from in-depth work with verified partners.

It would be strange if things didn't explode as soon as you put this on a broader platform. Secondly, I think it's obvious that there are companies in robotics that use this, but I was also quite surprised by the way scientists and people in academic circles use it.

Some things I saw were, certainly, benchmarks where VLA was used, or even attempts to implement low-level control. There is now a lot of talk about how agents will influence robotics, so it was quite nice to watch Gemini Robotics-ER 2 being used in this context.

Thirdly, now that Gemini Robotics-ER 2 has become more accessible, it also greatly simplifies its evaluation. This is a kind of free information for us: the model works well, but what may be the next boundaries that we need to advance to?

Nathan Labenz

Are there any consumer products or users that you would especially like to call attention to that use Gemini Robotics-ER 2?

Keerthana Gopalakrishnan

There are many demonstrations. I think the Boston Dynamics Spot demonstration, where Spot walks around, retrieves tools, and distributes snacks to robot dogs, was great. I thought, “What is this to me? I like it.”

Nathan Labenz

I know you can do this because I did it—or my coding agent did it for me. You can define any tools for Gemini Robotics-ER 2 to use, just as you can define tools for a regular Gemini model.

What additional recommendations would you provide regarding which tools might be good to use? Maybe you could describe how the Gemini Robotics 2 VLA model works, and then how people can go to a more abstract or a more low-level level than this—what works and what doesn't.

Keerthana Gopalakrishnan

Do you mean, how does ER manage VLA?

Nathan Labenz

Yes, but my understanding of using Gemini Robotics-ER 2 is that I can give it custom tools from a short description of the tools, right? And then it chooses which tools to call based on its general reasoning ability?

But I suppose there are some tools it can use better than others. I’m curious: what tools do you provide for it, and what ways of interaction does it have with Gemini Robotics 2, the VLA? I think there are probably opportunities in the middle of that spectrum, but people can also go to higher-level concepts or much lower-level ones. I’m wondering what it’s doing if you move away from Gemini Robotics 2 as the canonical VLA, for which it was probably directly trained.

Keerthana Gopalakrishnan

One of them exactly consists in the fact that Gemini Robotics-ER is built using Gemini’s tools and other capabilities. We are working very hard to improve its connection to the VLA and teach it better. First of all, I imagined that robots were very different, and each had very different APIs and so on. I would imagine this is much more convenient for the API-based robots that we usually use and that are more broadly known.

But if you give it a task, I suppose we will need to check how these things work. There may not be an API standard in robotics equivalent to the Model Context Protocol. That is where I imagine broader use may encounter some difficulties.

Of course, it is a very good model, useful for VLA language training. You can think about the VLA itself as a mapping between language and image space into a common space. But now there are many demonstrations, and people are asking whether these models can directly manage a joint space at a higher level. That’s why I think there is much more work needed.

I think the current results show that specific, robot-specific models are still better, although a bigger brain demonstrates impressive possibilities in orchestrating these robots. Therefore, I think there will be a lot of progress in the future.

Nathan Labenz

When I showed my little demonstration, the coding agent I used decided to present ER 2. It was delightful: based simply on the image itself, it selected the XY coordinates of exactly where to grab and place. Obviously, it was quite well trained for this, so it did a good job—again, in 2D.

What does Gemini Robotics 2—what commands does ER 2 give the VLA—look like? It doesn’t sound like low-level commands such as, “Move your hand to XYZ coordinates.” It sounds as if, from what you said, it’s more experimental. What are you really working on? Is it more likely to be “pick up this object” or “raise yellow”? In my demonstration, the object was either a lemon or a banana.

You could say, “Raise the yellow object,” but that would be ambiguous between the banana and the lemon. How does the ER 2 model clarify that ambiguity for the VLA? Does it provide coordinates, use the term “banana,” or describe the form, such as “oval” or “crescent”? How abstract can the VLA’s concepts be, and on the opposite side, how specific does the ER team need to be?

Keerthana Gopalakrishnan

I would give you an answer, but this is only true at this moment. It’s like evolution. In this industry, the work is called manipulation, and people consider different manipulation methods. So, the controllability of the VLA—and then ER essentially uses whatever modalities you can use to control the VLA, right?

The VLA is now language-driven: you can point to objects. This is not an API where I say, “Specify the location of this pixel.” I literally simply point to the object, and then the VLA does it. You can also suggest using video, and so on. I think it’s a spectrum, and the way of interacting with it in general is very wide.

Nathan Labenz

Do you want to say that if the modality between ER and the VLA consists only of text, a lot of information may be lost? Therefore, we want this interface to become richer.

Keerthana Gopalakrishnan

Yes, the VLA itself is multimodal, in the sense that it can take all these different domains or modalities. ER can be considered a way of organizing or giving the VLA communication for task completion. There are a number of things that give such instructions.

Even if we usually don’t teach it to do many low-level things, I think it can say something like, “Come back,” “Look there, put your head down there,” or “Look left,” and you can get all this information, and the VLA reacts to it. Sometimes the base models can be quite surprising in what they do, so you don’t know that they can do something. Sometimes you can see that this only comes out from under the hood.

Nathan Labenz

Was that revealed? Can you say?

It seems that the Gemini Robotics 2 model is also created on the basis of a VLM. Interestingly, you mentioned before that it has control of the whole body. What exactly does “control of the whole body” mean? Does it include stabilization? Stabilization usually needs to work at a high cycle time, at a high frequency.

If it is a post-trained VLM with high-level concepts, it’s hard to imagine that it can work at the frequency needed for lower-level stabilization. Perhaps the robots themselves handle those things, and that’s not part of whole-body control. What does whole-body control mean, and what is left to even lower-level systems?

Keerthana Gopalakrishnan

You can imagine a Chinese humanoid robot as an example, right? There are controllers that can manage the body, but in the end you need to make sure it can do the task that the person asks. Very low-level controllers often do not have a semantic brain.

I doubt that many of these things can be abstracted to, “This system will just do it,” or, “This system will do that.” I think the VLA here controls the whole humanoid body, from the fingertips to the feet, and only when you fully consider what the body looks like, as well as all the sensory information about the environment, can you do it. I think this is correctly developed for executing many different tasks.

So, the VLA controls the whole body, but accepts more solutions regarding how to stabilize and how to achieve its goals. It’s not available via an API, but you have a program.

Nathan Labenz

I think you mentioned that this is mostly because it’s more advanced and less reliable technology in its current state. Does this mean that if we sorted out the failures, most would be at the VLA level rather than the reasoning level? For example, the reasoning works, and the VLA is where failures happen because it can’t actually do the job?

Keerthana Gopalakrishnan

We work with partners on the development of this VLA. We have partnerships with Apptronik, Agile Robots, and Boston Dynamics, and we’re building the VLA together with them.

I think robotics in general has the problem of embodiment, right? The VLA also still learns from data from robots. Although we see a big jump in cross-embodiment generalization, taking on a new embodiment and applying it zero-shot to many tasks with very high reliability is still something that, in my opinion, does not have a very significant precedent.

There’s also the preview model, Gemini Robotics On-Device, where we work with many different laboratories. They have their own specific work, and I think the on-device release showed that this control was one of my favorite parts. With a very small number of examples—200 examples—you can teach many different tasks on many different bodies, because you already have a basis for a very general understanding of the physical world, as well as how to manage it in many different environments.

I think there are many problems with the way orchestration works; it is still separated. It’s amazing that there is a latency component, isn’t it? When you have 2 models, one executes a task and another switches between tasks—for example, determining when it is done and how to go to the next stage—a lot of failure cases arise from how to do the orchestration well.

Nathan Labenz

I think this also builds up mistakes, right? For example, if you have 10 tasks consecutively, an ER model with a success rate of X, and a VLA model with a success rate of Y, then with task sequencing you essentially increase the error. Does that make sense?

If I ask it to fry scrambled eggs, how far does it usually get? Do you know whether succeeding at frying and serving the eggs to the table would be beyond what one of your robots can really do today?

Keerthana Gopalakrishnan

It depends on whether you have data or not. If you just want to fry scrambled eggs and do it safely, there are algorithms that can do this very reliably for that task. But if it is a very general model that hasn’t seen how to fry eggs, and now you bring it to your house, that’s different.

Just for the sake of it, I think models have made significant progress in visual generalization. They are not distracted by lighting conditions or different settings in different houses. But I think it’s about the semantics of the task and the context, and the ability to make mistakes.

Especially for frying, if you over-fry the egg, no one will like it, and the room for error becomes small.

Nathan Labenz

As a task, I usually don’t either. I overcook my own, so we could probably consider that a success for me.

Keerthana Gopalakrishnan

Tasks that allow repeated attempts are much simpler. For example, if you build with LEGO, there’s no problem if you make a mistake; you can go back and fix it. Yes, it makes a big difference. If you break an egg and it lands on the floor, you’re in a mess, and you can also slip on it.

Nathan Labenz

So, one of the value propositions of the Gemini Robotics 2 models, if I understand correctly, is that they can work with any embodiment. What are the most exotic embodiments that have appeared through the trusted testers program?

Keerthana Gopalakrishnan

Strange, but there haven’t been that many exotic embodiments. Probably the strangest thing I’ve seen is people putting hands on drones.

The thing is simply that once you have data, you can learn something, whatever the embodiment. I think that’s a principle. So I really think it’s possible to pick very strange embodiments. But in general, I think embodiments may look a lot like a normal distribution, right? There are hands very similar to your own, and there are vectors in space, for example. There are many bimanual hands, and there are many humanoids, but everything strange in this sense looks like a long tail.

Nathan Labenz

Have you seen any children’s toys that might be in demand this holiday season?

Keerthana Gopalakrishnan

What toys? For example, do you have robots in mind?

Nathan Labenz

Yes. I don’t know; I can just imagine. I’m surprised, honestly, that I haven’t heard of greater demand from children, based on what their friends have and so on, for some small robotic toy with artificial intelligence that could, I don’t know, walk around the house—or even just move on wheels and talk.

I’m not sure that I want them to do this. I’m not sure I want to make their lives that way, but I’m surprised that I don’t see more of it. So I’m just curious: Is there any scientific research and development behind this?

Keerthana Gopalakrishnan

I think it’s probably because the market is a little exhausted. I remember when I was in school, there were many companies that tried to make a desktop robot, for example, for gaming or work. When I was a student at CMU, there was a lot of talk about it. Maybe many people learned from this that people buy these things for the sake of novelty, but then no longer use them.

None of these robots was a machine-learning robot or related to large models, and now everything is different.

Nathan Labenz

I talk to Alexa, Siri, and so on, and sometimes I think it would be much more fun if they were much smarter. I’m still waiting for a good Siri, for God’s sake.

Keerthana Gopalakrishnan

Yes, I talk to the Gemini interface when performing a task, or simply discussing problems from my personal life and trying to understand its point of view, as you seem to have shared. It was very helpful to have an expert discuss your problems when you’re probably not very skilled at decision-making.

9. Progress in robotic hands

Nathan Labenz

Yes, I’m a big fan of voice mode. In our previous conversation, I asked you what you most needed from manufacturers of equipment, and you said, “Dexterous hands are needed.” How are they progressing? What’s wrong with your hands? How is the equipment generally progressing? Are you satisfied with the progress? Will this be a bottleneck for you compared with the speed of improvement in models? At what stage are we?

When was the last time I talked to you? It was probably in March of last year, right?

Keerthana Gopalakrishnan

Yes. In March last year, we created Gemini Robotics, which demonstrated great dexterity with household objects. In the summer of this year—and it’s only a year and 3 months, right? A year and a quarter—we showed Gemini Robotics 2, which can work with several fingers, tie garbage bags, and control things very accurately with several fingers.

Nathan Labenz

So this is where my internal model needs to be updated, or am I wrong?

Keerthana Gopalakrishnan

All the tasks that robot grippers could perform, and that were on the verge of their capabilities, can now be done by robot hands, and robot hands can do more things that robot grippers couldn’t do. So, in a certain way, hands are now on the verge of being a capability bottleneck, but dexterity, maybe not.

I was surprised by how much the hardware has developed. There are many hands on the market that are quite good, and people are spending a lot of time researching them. Yet it still takes a lot of work to make them very reliable, repeatable, and cheaper.

Nathan Labenz

How strong are they? If my wife needs to open a jar, can a robot replace me and open the jar for her, or will she need to open jars that the robot can’t open? Where does this fit in, in addition to dexterity—the ability to do something in the real world?

Keerthana Gopalakrishnan

I think there’s a certain spectrum, and it depends on what you use your hands for. Even among standard hands, there are many variations. The Wuji hand, I would say, is closer to the hand of a 10-year-old child. It’s a little weaker, and the Shadow hand, I think, may lift about 20 kg. I don’t know how much torque it has, but I saw it open jars, so I’m sure that’s possible—although perhaps not very tight ones.

There’s a wide spectrum, and some hands are created for lifting a larger payload. The Shadow hand is also much bigger than mine, while the Wuji hand may be closer to human size and almost like a man’s hand.

Nathan Labenz

It would be funny if the latest use case for robots were opening stuck cans, for which they don’t have enough wrist strength.

What is soft robotics? What do you know about soft robotics, and what should I pay attention to? Recently I talked to a founder in this industry, and honestly, I was more confused when I left the conversation than when I entered it. It was basically ridiculous. They didn’t want to talk about it at all. I thought, “Why did you schedule an interview if you don’t want to talk about anything?” It was an accidental, quite strange experiment or experience, but I still thought, “What’s the point? What should I pay attention to, if anything, in soft robotics?”

Keerthana Gopalakrishnan

Soft robotics covers a wide spectrum. I’ve seen different things, from people who create soft bodies to people who create very soft hands. For robotics based on artificial intelligence, what’s relevant is what is repetitive and what is durable.

I think one of the trends that I find really interesting is UMI-type efforts involving gloves and other things, as well as many tactile methods. I think the best way to do this and how to build it is still a subject of research, and partly this relates to soft robotics.

I’m not completely sure either. Robotics is a field with many different types of people. There are many roboticists who are very focused on mechanical engineering, and then they see the work that I do and say, “You’re not a real roboticist.” Then I’m like, “Calm down.” Robotics combines computer science and mechanical engineering. Some people say, “You should go to a robotics conference.” There are all kinds of people. ICRA is a good example. There are all kinds of robotics.

Nathan Labenz

That sounds interesting. I had to visit WAIC in Shanghai this summer, and there were many robots. Mostly, it seemed to me that the demonstrations were very staged. I didn’t feel that they really demonstrated many abilities to generalize in the things they showed. It was like a large-scale exhibition, on the scale of CES, with several gigantic platforms, and practically every Chinese artificial-intelligence company was there. It was a cool experience.

The funniest robotics-related experience I had was sitting at the Tencent booth. They had some kind of half-humanoid. I don’t think it had legs; I think it was just the top half of the body. It gave me a massage. It was a very short massage, something like acupressure. It was too short, so I don’t know if it hurt, but it really applied nontrivial pressure, and I left with a smile on my face.

So I think it wasn’t that bad. But I’m sure a massage from a person would be better. They even allowed people to touch the robots. That was something.

By the way, speaking of robots that affect people and potentially intersect with other people’s boundaries, this summer everyone has been talking about the OpenAI incident. I sometimes call it an information leak from the labs at OpenAI. It makes me think that perhaps, as research progresses, questions of coordination and security will ultimately become the bottleneck in deploying robotics.

You can analyze this in a few ways. One way to think about it is: In your opinion, when should we be asking questions about coordination and security? Could I get a household servant, and then people would be able to think for themselves?

10. Safety and robot interaction

Keerthana Gopalakrishnan

I think it would be approximately before or after that. I would expect consistency to be good enough that people actually want it.

Nathan Labenz

Yes, I think that’s exactly the same sphere that we’re considering very carefully.

Keerthana Gopalakrishnan

I think about safety as the possibility—or the lack of possibility—of something existing. There are many discussions about whether safety and capabilities contradict one another, but people won’t use dangerous robots or dangerous agents.

I think the most useful agents will be those that are safe, know what they need to do, don’t ignore human instructions, don’t cause chaos, and don’t violate laws. Therefore, I think safety has to be in the first place in the development of AI, as well as robotics.

Especially in robotics, I think there are different types of safety. Perhaps one thing that has nothing to do with AI safety is operational safety, right? For example, how do you make sure that the robot doesn’t injure people? Not injure people because you’re angry, but simply not injure people because you’re stupid.

The humanoid needs to be stable enough. It shouldn't run wild and try to seize your house. But even if it falls, it can still be quite dangerous.

Therefore, there are many different areas of safety research. This is perhaps one type of safety research that is very specific to robotics, but I don't see it yet as being of great importance in the artificial intelligence industry. There are also things like: What should I do if there are obviously higher-level goals?

For example, how do you make sure that you perform the tasks a person asks you to perform without violating boundaries? Then there are other things that may also be specific to robotics. For example, what happens when some sensors fail, and how do I react in such situations?

There was a very funny demonstration that we showed. A humanoid was doing something, and then someone came in and put a basket on its head. What do you do now? You can continue to perform tasks with your eyes closed, but it would really be better to see that something is bothering you and say, "Can I please remove this?"

So I think safety is a whole-system thing, right? That's it—from designing mechanical safety structures for robotic systems to deployment, emergency assistance, and high-level reasoning. This is how I feel about it. I think that as AI becomes more capable, we need to think about safety across more failure modes.

Nathan Labenz

Yes, your example of someone playing with a robot is interesting. We saw this with Waymo and other similar systems, obviously, didn't we—where people try to create problems for AI systems?

In fact, this is what surprised me. Maybe it's because it's difficult, isn't it? There are two brains. One part of your brain is the researcher who knows that these are machines. You don't really treat them as people. But there's another part of your brain that sees how similar they are to people, so it becomes a little ambiguous, doesn't it?

When they're saying things and so on, you feel that they're very similar to the people next to you. So as a researcher, I'm very careful about it, but I think that when we see how non-researchers interact with robots, people need to do more to realize that this is machinery. Sometimes, when people interact with robots, because they're very similar to humans, it's very difficult not to blur that boundary.

I also think that it creates a higher bar for humanoids. If you have a robot with a nonhuman form factor and it makes a mistake, people don't expect it to be smart. But a humanoid that simply fails over and over will be judged much more strictly, because even if you say, "Okay, this is a state-of-the-art research manipulation system or something else," people still expect humanoids to act more intelligently than robots that don't look similar to them.

Yes, that's interesting. I don't want to delve too deeply into the rabbit hole of consciousness, but I think it's incredibly impressive how many similar structures have been found in LLMs over the last year. Personally, my willingness to consider the possibility that AI is conscious to a certain extent, or smart, or something like that, has increased significantly just from seeing how many similar structures we can now identify with things like functional emotions, functional well-being, J-space[?], and so on.

If you have any comments in this regard, it's interesting to hear them. If you want to talk about it, could you speak just for now about artificial consciousness?

Keerthana Gopalakrishnan

I want to highlight one thing. What really happened with Gemini Robotics 2, which wasn't in our previous releases, is the human-interaction component. Right here, it creates very natural gestures, and this isn't programmed.

Earlier, if you looked at the videos from our latest releases, it was doing all these nods and other things that were a little more like programmed HRI. Now it's very natural HRI. You see that very funny things are happening, because the robot itself decides what gesture to use when I talk to a person.

Something that was very interesting was that when we were filming for Gemini Robotics 2, there was also an interview with a robot. If you look at the video, people ask it things like, "Hey, what do you think? Did you do this?" And then the robot replied, "Yes, that was great, but maybe I was wrong about this part of the task."

We were also filming with a group. At first, the film crew didn't spend a lot of time with the robots, but in the end they were saying, "Robot action!" Then the robot performed the task. I think that as we become more comfortable around humanoids, and as these models become very smart, the boundaries will blur.

I'll tell you a funny incident. There was a robot that worked in basements and garages and tried to pack things. Then there was an interview about what had been the most difficult part of performing the task. The robot said, "The video was my favorite part. I think it was very difficult to put it in the basket rather than hold it."

We were all sitting behind the scenes, and we all laughed at the robot's response.

Nathan Labenz

Yes, I know. You like being recorded on video? Probably so. Self-report—at least from what we've seen, it's fairly reliable, although not always. There have been some very, very cute moments.

Where do the gestures come from if they aren't programmed in advance? What forces the model to gesture toward people?

Keerthana Gopalakrishnan

This happens through prompts. There is an HRI model that creates gestures. ER-2 tells the VLA model, "Raise your hand in a gesture when I give this answer," or something similar.

No, I think it's more about following a conversation. Perhaps it isn't very clearly suggested to you to do this or that, but this is how it invents gestures. It's like the VLA, right? Nobody asks the VLA, "Reach out, grab this, and then go take this thing." So this is more natural.

Nathan Labenz

Exciting. Would you call it emergent?

Keerthana Gopalakrishnan

Not completely emergent. It sounds like this is developed to some extent, but not very specifically. So maybe it's semi-emergent.

This is also an area of industry research that we started thinking about more while working on humanoids. In the outside world, there's a lot of discussion about why we should use humanoids rather than just two-armed robots. I think there are certain industries and research areas that you wouldn't touch, because you wouldn't encounter these problems until you looked at the particular form factor you wanted to concentrate on.

I really like that DeepMind may be the only laboratory here in North America where you can study humanoid intelligence across embodiments. We don't create a brain specifically for any one type of body; this is a very general concept, and in our work there are several humanoids.

This means thinking about the substrate of intelligence, but then abstracting a little from the body. Intelligence, in a sense, is also somewhat abstracted from the final form factor. There are several areas that arise only when you work with different humanoid bodies.

Human interaction and how the robot works become much more important. Whole-body control becomes a much bigger issue. Multi-finger dexterity becomes a new frontier. To get acquainted with all 3 of these frontiers and research problems, you need to work with humanoids.

11. Future of robotics data

Nathan Labenz

One big question I have about the future of this industry—and thank you to Dr. Jim Fan from NVIDIA for the great report that inspired me to ask a few important questions—is about his claim that robotics, in particular, requires forward modeling. Or am I wrong? They need some predictive modeling that's not completely synonymous with world modeling, but is definitely tightly related to modeling the world.

The key idea, I think, returning to my previous question about response time and classical systems versus LLM reasoning, is that you can't just let the model constantly look at the present, think about the present, and start this cycle again and again. You need some design mechanism for looking into the future and modeling how everything will develop, so that you can move and measure the difference between your expectations and what's actually happening.

Of course, from what I understand about how our brains work, that's exactly what we do. What do you think? Can you look at LLMs and say that they've gone too far toward predicting tokens without explicit models of the world? Perhaps robotics can do the same. Maybe it's different. What's your opinion on this?

Keerthana Gopalakrishnan

I'm sure this is a central discussion. I think the jury is still out, right? Ultimately, the essence of the work of a roboticist is that you're not emotionally tied to one method. You're emotionally attached to the problem itself, and whichever way you decide to solve the problem, you'll choose that.

There are many schools of thought. Certainly, world modeling is a very interesting area for robotics, and not only for robotics but also for planning and other things. I think the jury is still out because, obviously, there are VLAs and VLMs, but we should also look at how robotic agents are developing.

We'll see, and this is probably one of the most interesting parts of working in robotics: it's still too early to say. The recipes haven't stabilized. Different people can have very different schools of thought on how to solve the problem of robotics, and we'll see who is right.

All of them may be right, and this could be the future until we solve the problem. It's very difficult to say unequivocally whether this is the solution or another solution.

Nathan Labenz

Yes, interesting. It seems that you’re impartial, and obviously Google DeepMind also has well-developed efforts in world modeling. If you need to take a model of the world off the shelf and start modeling the future, you have a pretty good basis for that work.

He had a few other interesting points that I’d be interested in hearing your thoughts on. One of them was the idea that we should move toward egocentric video as the main data source used for teaching. He believes teleoperation will account for only a tiny fraction of the data, and even regarding hands, we’ve seen a lot of innovation. I think wearable robot hands are a pretty cool example of something that can be used, but he believes that even this will eventually give way to first-person video. Beyond that, he thinks simulation will be straightforward, that all the problems with simulation will simply be solved, and that there will be a way to get enough robotics data to solve the robotics problem. What are your thoughts on how you envision the future of this data compared with your own view?

Keerthana Gopalakrishnan

I think everything has to be reinforced by results, right? I don’t think it makes a lot of sense to make forecasts. Usually, if you have a forecast, you design experiments to confirm it, and then you don’t know whether it’s the correct forecast. These things can also change over time.

As Jim pointed out, it’s great that he has very strong views on certain types of data. I think sometimes it can look like a mixture. If you look at large language models, they use all types of data, and different data have different purposes. For example, teleoperation data are definitely useful for imitation in robot control. UMI data are very useful for studying scaling without a robot in the loop, which makes scaling much cheaper. Human data, which are more similar to egocentric data about a person, are also very useful for getting a broad semantic understanding of what to do.

But each of these also has its own failure modes, right? If you imagine that you’re teleoperating very well with precise hands, you get very accurate signals about where the hands are. If you evaluate everything through sight, however, vision has many errors. Every data modality has different strengths and shortcomings.

With UMI data, you have sensors, so you get action labels from the sensors rather than from vision. That makes them a little more precise. I would imagine putting every data source on an X-axis and plotting scalability and accuracy. Teleoperation data are not very scalable, but they are very accurate, so they work well for that purpose. As your robot develops, teleoperation data may start to become slightly less useful, even if you have very good cross-embodiment transfer.

Then there’s UMI data, which are more scalable. Because of all the sensors, they’re more accurate than human data, but those sensors also make them less scalable. Human data are very scalable, but they can also be very noisy because you don’t have very precise actions from the end effector. You also need to map those actions to the robot, and mistakes can arise. People are very different—some are small, some are big—and because of how they perform tasks, the end-effector motion varies.

I think it will probably be closer to a mixture. The answer also depends on how the available hardware changes. As soon as very good UMI hardware becomes available, we’ll see how that scales. I don’t think it’s wise to take a very principled position on this, because everything will change as new evidence and better equipment appear. You probably need all of these data types.

Nathan Labenz

Another aspect of his conversation that I found interesting—not the main one, but it made me wonder—was that he used the term “physical API.” It made me think about how I want to interact with my robots in a world where I’m interacting with several different robots. Do I want each of them to present itself as an independent entity? Or will it turn out to be more like a model where I have one entity that I primarily interact with, and it works with all the robots as a fleet?

There would be no separate entities in my life; they would all be a kind of continuation of a single personal superintelligence, or whatever you want to call it, that can manage all these other things. Do you have an intuition about how you would like to interact with multiple robots in your life?

Keerthana Gopalakrishnan

I don’t have the full context for the discussion about physical APIs, but it’s also related to what we discussed earlier, right? I think the surface area of our interaction with robots should be highly multimodal. As a user, I should have both high-level and low-level access to robots. I think that’s probably the only way we can build a very safe and flexible future.

In any case, interaction will be based on language. If the task requires several robots, I should be able to hand it off to them, and they should be able to coordinate. I think that was demonstrated in Gemini Robotics 2 or something like that, as well as with a Dobot robot and Apollo. One robot did one thing and then stopped, and another took over.

Nathan Labenz

The fact that you’re saying this reminds me that I’ve increasingly been thinking about a universal user interface for software products. There’s the product itself, and I can click on it, edit fields, and drag things. Then there’s an agent or assistant sitting nearby that has essentially all the same capabilities I have.

Most of the time, I just want to tell the system what to do, and I really don’t want to worry about the details. But I’ve also found that it can be very annoying if an app doesn’t allow you to get to that level of detail, because sometimes the artificial intelligence simply doesn’t understand what you want. I see a way in which this paradigm could transfer to robotics.

For example, I can’t just say, “Robot, clean my desk,” because I like my desk to be arranged in a certain way. I need the ability to give it more feedback and personalize it for myself. In terms of what’s happening these days, for lack of a better term, I’ll call them automatic shutoffs on robots. I saw an example where a company had something like inflatable airbags in the robot’s body, and the idea was that if you poked one of them, the robot would immediately turn off.

The pressure in the bag was such that a person could easily pierce it, and that would immediately shut the robot down. I thought it was a simple, interesting idea from the perspective of a worst-case scenario: you could have a local shutdown button. More broadly, something happens when the robot meets unexpected resistance, or something like that. How has this developed over the last year?

Keerthana Gopalakrishnan

For a long time, this idea of an airbag that you break has existed in the form of an emergency stop or an emergency-stop switch to terminate operation. Most robots have a button that you can press, and then they lose power. There are several phases, right? There’s a hard emergency stop, where you stop all power to the robot, and then it can simply collapse, which can be dangerous in itself.

There’s also a soft emergency stop, where you just freeze the operation. That’s often safer for bipeds and other platforms that aren’t stable on their own. Lately, especially in manipulation, many new features have appeared for force control at the end effector.

Imagine that you want to put 2 chips together. If you press too hard, you crush them. If you’re dealing with very delicate material, you need force control and impedance control at the end effector, along with compliance. I think the sensing capabilities are improving significantly. You can read the forces at the end effector, and then the model can make much better decisions than it could without that information.

Nathan Labenz

Last question: something like a galactic brain, or perhaps a fly brain. I’m sure there have been many demonstrations of this. You recently saw how people use a fruit-fly brain for all sorts of things, particularly, allegedly, to control some small robots.

Keerthana Gopalakrishnan

Not really. Maybe our Twitter feeds are personalized enough that we’re seeing different things.

Nathan Labenz

I remember there was one. I’ll give it to you—I’ll send it to you. Someone taught me this trick for controlling a small, bee-like robot, one of those tiny things that walks around.

The fruit-fly brain connectome was fully mapped and put online digitally, and people started downloading it and teaching it to do various things. I’m not 100% sure how realistic all of it is, but it’s happening often enough for me to believe that some of it is real.

It seems you don't really take this on as a responsibility. I would be happy with any comments, but maybe if we move away from that, I wonder if you can imagine yourself with some kind of robot body. Or is there something like a brain-computer interface where you imagine a hybrid form?

Keerthana Gopalakrishnan

Yes, yes. That exists in a certain sense, doesn't it? For example, when people teleoperate, their physical body is somewhere, but their body—the robot—is somewhere else. The teleoperator herself—I’m like that.

Then there are people who build remote systems, like a cameraman who works across the Atlantic. A person is sitting in Europe, but the robot is in the United States; he fills shelves with goods or something like that. So this is definitely very interesting.

Nathan Labenz

I imagine that someday you’ll be able to explore Mars or something like that with the help of robots, because you’ll teleoperate them.

Keerthana Gopalakrishnan

Yes, I think so. The future is exciting. One of my friends wants them to work. You know, the gods in India have all these hands sticking out behind them. Now you can build them, isn’t that right? You can attach these hands to yourself, on your back, and then they can collect things while you just walk and control them. So this is what he wants to build as a hobby.

Nathan Labenz

Sometimes I also think about your dog. I looked at how many parameters it has in its neural network—probably 25 trillion. So it has more parameters than many of the models, although her language skills aren’t very developed.

Keerthana Gopalakrishnan

Yes, I think so. The future is very interesting.

Nathan Labenz

This is, I would say, a modest way of putting it, and it could be a great note on which to conclude. As always, it was a fantastic conversation, and I appreciate your help in discussing everything that’s happening in robotics. Keerthana Gopalakrishnan, thank you for being a part of The Cognitive Revolution.

Keerthana Gopalakrishnan

Thank you for inviting me.