Robotics AI
Reading time 5 min readGemini Robotics

Gemini Robotics 2 and ER 2 connect vision, language, planning and robot action

Google DeepMind's current Gemini Robotics 2 family includes a vision-language-action model for robot control and Gemini Robotics-ER 2 for embodied reasoning and multi-step physical planning.

By TechniaHQRobot

Google DeepMind now presents Gemini Robotics 2, Gemini Robotics-ER 2 and an on-device variant as its current robot-model family.

Gemini Robotics 2 is the action model

Google DeepMind describes Gemini Robotics 2 as a vision-language-action model that can control robot motion from visual input and natural-language instructions. The company shows it across different robot embodiments, including whole-body humanoid control.

The robot body still sets physical limits. Joint range, gripper geometry, camera placement, payload and safety controllers affect whether a policy can execute a planned action.

Gemini Robotics-ER 2 handles embodied reasoning

Gemini Robotics-ER 2 is positioned as a higher-level model for spatial reasoning, multi-step planning and physical-world understanding. Its model card lists text, image, video and audio inputs with text output.

ER 2 can pass plans or instructions to lower-level robot controllers. Google also documents concurrent reasoning during action and collaboration across more than one robot.

Robot-specific testing remains necessary

The ER 2 model card warns against using the model as the sole controller in safety-critical production settings. Physical execution still needs robot-specific validation, motion limits, collision handling and recovery behavior.

A useful evaluation records task success, intervention, contact errors, planning latency and failure recovery on the exact robot and environment.

Current evidence boundary

Google's published demonstrations show manipulation and whole-body tasks, but they do not provide a universal autonomy score that transfers across homes, warehouses and factories.

Claims should stay attached to the named model, robot, task and test conditions.

By @techniahqrobot

About the publication · Sources and editorial policy · Report a correction

Evidence reviewReviewed 2026-07-23

Gemini Robotics evidence and model boundaries

Gemini Robotics is presented in the research report as a vision-language-action model built on Gemini 2.0, with a related embodied-reasoning model for spatial and temporal tasks. The paper reports manipulation across defined robots, tasks and test conditions. These results should not be generalized to every embodiment or treated as proof of unrestricted real-world autonomy.

Verified context

  • The report describes direct robot control through a vision-language-action model.
  • Gemini Robotics-ER covers embodied reasoning outputs such as object detection, pointing, trajectory and grasp prediction.
  • The experiments include fine-tuning and adaptation to specific capabilities and robot embodiments.

What the available evidence does not prove

  • The paper does not establish universal performance across untested robots.
  • Reasoning predictions still require a safe control system and physical validation.

Sources