Google DeepMind now presents Gemini Robotics 2, Gemini Robotics-ER 2 and an on-device variant as its current robot-model family.
Gemini Robotics 2 is the action model
Google DeepMind describes Gemini Robotics 2 as a vision-language-action model that can control robot motion from visual input and natural-language instructions. The company shows it across different robot embodiments, including whole-body humanoid control.
The robot body still sets physical limits. Joint range, gripper geometry, camera placement, payload and safety controllers affect whether a policy can execute a planned action.
Gemini Robotics-ER 2 handles embodied reasoning
Gemini Robotics-ER 2 is positioned as a higher-level model for spatial reasoning, multi-step planning and physical-world understanding. Its model card lists text, image, video and audio inputs with text output.
ER 2 can pass plans or instructions to lower-level robot controllers. Google also documents concurrent reasoning during action and collaboration across more than one robot.
Robot-specific testing remains necessary
The ER 2 model card warns against using the model as the sole controller in safety-critical production settings. Physical execution still needs robot-specific validation, motion limits, collision handling and recovery behavior.
A useful evaluation records task success, intervention, contact errors, planning latency and failure recovery on the exact robot and environment.
Current evidence boundary
Google's published demonstrations show manipulation and whole-body tasks, but they do not provide a universal autonomy score that transfers across homes, warehouses and factories.
Claims should stay attached to the named model, robot, task and test conditions.
By @techniahqrobot
About the publication · Sources and editorial policy · Report a correction