
Google DeepMind Gemini Robotics 2 extends VLA control from tabletop manipulation to full humanoid motion and dexterous hands. Here is what the demos prove and what remains difficult.
Introduction
Walking to an object and manipulating it is a different problem from manipulating an object already placed in front of a stationary robot. Gemini Robotics 2 matters because Google DeepMind has moved its VLA stack into that larger control loop: the same task can require stepping, crouching, reaching, balance and hand control before the grasp even begins.
DeepMind introduced Gemini Robotics 2 on July 30, 2026. Its published evaluation spans Apptronik Apollo 2 with different hands and a Franka Duo with parallel grippers. The result is important, but the data also shows the limit clearly: whole-body and gripper tasks can be strong while multifinger dexterity remains substantially harder.
Research date: August 12, 2026. Information verified from official sources available as of August 12, 2026.
Direct answer
Walking to an object and manipulating it is a different problem from manipulating an object already placed in front of a stationary robot. Gemini Robotics 2 matters because Google DeepMind has moved its VLA stack into that larger control loop: the same task can require stepping, crouching, reaching, balance and hand control before the grasp even begins.
Key findings
- Gemini Robotics 2 is a VLA model that converts multimodal inputs into motor actions and now controls complete humanoid bodies rather than only upper-body tabletop behavior.
- DeepMind demonstrates one model checkpoint across Apollo 2 configurations and Franka Duo, providing evidence of cross-embodiment reuse within its supported setups.
- Reported Apollo tasks include pick-up from table 68.4%, floor 45.7% and shelf 76.3% with Inspire hands.
- Reported Franka Duo results include 74.2% general pick-and-place, 78.9% tool kitting and 89.6% precise insertion.
- The five-finger 22-DoF SharpaWave hand can attempt knot tying and ziplock sealing, but DeepMind explicitly says multifinger manipulation remains challenging.
What to measure before treating the claim as deployment evidence
A reusable evidence checklist applied throughout this article.
| Signal | Useful evidence | Common mistake |
|---|---|---|
| Capability | Repeated task success with defined trials and resets | Judging one edited demonstration |
| Autonomy | Human intervention and teleoperation disclosed | Calling scripted or supervised behavior autonomous |
| Generalization | Unseen variable and adaptation budget stated | Using zero-shot without defining what was unseen |
| Reliability | Long runs, recovery and failure logs | Reporting only peak performance |
| Deployment | Customer workflow, uptime and support burden | Equating hardware shipment with productive use |
Not every row applies equally to research papers and public-market analysis; the article specifies the relevant evidence.
What changed from the previous Gemini Robotics generation
Earlier Gemini Robotics demonstrations concentrated heavily on upper-body and tabletop tasks. Gemini Robotics 2 expands the action space to the entire humanoid. A command can now require the robot to walk toward a shelf, change its center of mass, bend or crouch, position an arm and manipulate the target while maintaining balance.
That coupling is technically important. Locomotion and manipulation cannot always be solved as independent modules because reaching changes the body's stability and the body pose changes which grasps are possible. A useful whole-body policy has to reason about the task while respecting those coupled constraints.
Apollo 2 turns semantic instructions into body motion
DeepMind's example asks Apptronik Apollo 2 to move a watering can to a green bin on a bottom shelf. The robot walks to the object, grasps it, takes several steps and places it at the destination. The point is not the watering can; it is the transition from language to a sequence that spans navigation and manipulation.
DeepMind also notes that movement speed still needs improvement. That caveat matters commercially. A task that succeeds at low speed in a lab may not meet production cycle time, and faster whole-body motion narrows the control margin around balance and contact.
The published numbers show where the difficulty moves
With Apollo 2 using Inspire hands, DeepMind reports 68.4% for picking from a table, 45.7% from the floor and 76.3% from a shelf. The floor result is a useful reminder that a seemingly small geometric change can force a different stance, reach envelope and balance strategy.
On Franka Duo, two-finger grippers perform better on several precision tasks: DeepMind reports 74.2% for general pick-and-place, 78.9% for diverse tool kitting and 89.6% for precise insertion. These are developer-reported averages under DeepMind's task protocols, not a universal robot ranking.
Five fingers create a much harder control problem
Gemini Robotics 2 controls a five-finger SharpaWave hand with 22 degrees of freedom on Apollo 2. DeepMind shows knot tying and sealing a ziplock bag. Each finger adds possible contacts, and contact changes what the next joint motion should be. The controller has to handle underactuated objects, occlusion and several acceptable grasp configurations.
DeepMind explicitly states that multifinger dexterous manipulation remains challenging. That distinction is essential. A model can understand the language instruction and still fail because fingertip pose, friction, force or timing is wrong by a small amount.
Cross-embodiment does not mean arbitrary robot control
DeepMind shows the same model checkpoint controlling Apollo 2 with different end effectors and a Franka Duo. That is meaningful evidence that a shared model can span distinct action interfaces when the training and system design account for them.
It should not be read as plug-and-play control of any unseen robot. Kinematics, joint limits, cameras, control rates and safety interfaces still differ. DeepMind's On-Device 2 work separately addresses adaptation to new bi-arm embodiments, which reinforces the point that embodiment transfer usually needs an explicit adaptation mechanism.
What would make Gemini Robotics 2 deployment-grade
The next evidence should include longer runs, disturbance recovery, intervention rates, control latency, repeated trials on unseen layouts and performance under human proximity. Whole-body capability increases the value of the robot but also increases the space of unsafe or unstable motions.
Gemini Robotics 2 is best read as a major expansion of the learned action space. The remaining problem is reliability: whether these abilities can run fast, repeatedly and safely enough to beat a more specialized automation stack on real cost and uptime.
Limitations and missing information
- Performance figures are reported by Google DeepMind under its own evaluation protocols.
- The VLA model is available to early-access partners rather than as an unrestricted public checkpoint.
- Success on selected tasks does not establish long-duration autonomous operation in factories or homes.
- Cross-embodiment evidence covers supported configurations and should not be generalized to arbitrary robot hardware.
Conclusion
Gemini Robotics 2 closes an important gap between semantic intelligence and whole-body action. The model now has to solve the geometry of the entire robot, not only the hand in front of a table.
The published results are strongest for whole-body pick-and-place and parallel-gripper dexterity. Five-finger manipulation remains the sharper frontier. That is where vision, tactile information, force control and high-frequency action will increasingly determine whether a humanoid looks capable or actually becomes useful.
Frequently asked questions
What is Gemini Robotics 2?
It is Google DeepMind’s 2026 vision-language-action model for robot control, including full humanoid bodies and bi-arm robots.
Which humanoid is shown with Gemini Robotics 2?
DeepMind demonstrates Apptronik Apollo 2 with multiple hand configurations.
Can it control dexterous hands?
DeepMind demonstrates a five-finger SharpaWave hand with 22 degrees of freedom, including knot tying and ziplock sealing, while noting that multifinger manipulation is still challenging.
Is Gemini Robotics 2 open source?
No equivalent unrestricted public checkpoint is offered. DeepMind says the VLA and On-Device models are available to early-access partners.
Sources and methodology
Research was checked on August 12, 2026. Current-company claims use official company or government material where available, while financing and listing details are cross-checked with Reuters.
Research-paper performance numbers are attributed to the authors and are not treated as independent validation. Benchmarks with different robots, tasks, resets or success definitions are not ranked as if they were directly comparable.
Related TechniaHQRobot guides
Official image recommendations
Use the exact robot and generation named below. Confirm reuse rights with the source owner before publication or social distribution.
Fact-check report
Verified:
Share this article
Share the current TechniaHQRobot article page.
Continue reading
Open the latest robotics reporting, Physical AI analysis and hardware notes.