NVIDIA's current GR00T platform page lists GR00T 1.7 alongside teleoperation, synthetic-data and deployment tools for humanoid and general robot learning.
What GR00T 1.7 receives and produces
NVIDIA describes GR00T 1.7 as a vision-language-action model that takes camera video, natural-language instructions and robot proprioceptive state as inputs. It returns chunks of robot actions for the target embodiment.
The model sits inside a larger development stack that includes real teleoperation data, synthetic data, simulation and deployment middleware.
The architecture combines reasoning and continuous action
The earlier GR00T N1 technical release described a dual-system architecture with a vision-language model for reasoning and a diffusion transformer for continuous action generation.
The two stages operate at different timescales: scene and instruction reasoning first, followed by continuous action generation for the robot.
Training data spans real and simulated robot experience
NVIDIA documents training pipelines that combine teleoperated robot trajectories, synthetic data and other visual sources. Its July 2026 GR00T 1.7 material reports roughly 32,000 hours of real robot data and about 8,000 hours of simulated data.
Those figures describe NVIDIA's training corpus. They do not establish the same task success on every robot body because camera placement, hands, joint limits and control frequency differ.
Provider benchmark results need their test conditions
For GR00T N1, NVIDIA reported a 76.8 percent average success rate in its real-world GR-1 evaluation when using the full training-data mixture. That number belongs to NVIDIA's test setup and model version.
A deployment team should repeat target tasks on its own robot and record success rate, intervention rate, collision recovery and latency before using a provider benchmark as an operating estimate.
By @techniahqrobot
About the publication · Sources and editorial policy · Report a correction