HOURS ≠ TRAJECTORIES ≠ SIMULATION

Physical AI Training Data: What the Counts Measure

Compare six reported robot-learning data budgets by training stage, unit, included subsets and access limits. Counts are reported by the authors, not independently recounted.

6 retained recordsSources checked September 29, 2026Primary-source fields preserved per record
6 of 6 records shownFilters are local and do not create crawlable URLs

LingBot-VLA 2.0 pretraining mix

Official repository
Dataset / version
LingBot-VLA 2.0 pretraining mix
Training stage
Pretraining mixture
Author-reported scale
~50,000 robot hours + ~10,000 egocentric human-video hours
Inputs / measurements
Vision, language, robot state/action; human egocentric video in pretraining mix
Why this is not directly comparable
Human video and robot-action data are distinct components. The 60,000-hour total does not mean 60,000 robot-control hours.
Available artifacts / access
Repository links code and model artifacts. It does not present the complete proprietary pretraining corpus as a single downloadable dataset.
Source checked
2026-09-29

OpenVLA pretraining data

Official project
Dataset / version
OpenVLA pretraining data
Training stage
Pretraining mixture
Author-reported scale
970,000 robot episodes from Open X-Embodiment
Inputs / measurements
RGB, language and robot actions/state depending on source dataset
Why this is not directly comparable
Episode duration varies and no comparable hour total is given. Some underlying datasets also appear in other mixtures.
Available artifacts / access
Model/code are released; underlying Open X-Embodiment datasets have separate access and licensing terms.
Source checked
2026-09-29

RDT-1B multi-robot pretraining collection

Official project
Dataset / version
RDT-1B multi-robot pretraining collection
Training stage
Pretraining and fine-tuning
Author-reported scale
1M+ pretraining episodes from 46 datasets; separately, 6K+ ALOHA fine-tuning episodes
Inputs / measurements
Language, up to three RGB views, robot actions/state
Why this is not directly comparable
The 46-dataset mix and the authors' ALOHA collection describe separate stages. Shared dataset families prevent adding this budget to OpenVLA's.
Available artifacts / access
Code, weights and the authors' ALOHA fine-tuning dataset are linked. The model card's MIT statement is not a blanket licence for all 46 source datasets.
Source checked
2026-09-29

RoboMIND 2.0

Research paper
Dataset / version
RoboMIND 2.0
Training stage
Dataset release
Author-reported scale
310K+ real-world trajectories; an additional 20K simulated trajectories
Inputs / measurements
Multimodal robot data; includes 12K tactile-enhanced and 20K mobile-manipulation trajectories
Why this is not directly comparable
The 12K tactile and 20K mobile-manipulation trajectories are subsets of the real-world data, not extra totals. Six embodiments and 739 tasks describe coverage.
Available artifacts / access
The paper describes the dataset. Complete file availability and dataset licence were not established by this audit.
Source checked
2026-09-29

TurboVLA AgileX Piper task demonstrations

Research result
Dataset / version
TurboVLA AgileX Piper task demonstrations
Training stage
Real-robot fine-tuning
Author-reported scale
65 teleoperated demonstrations per task × 4 tasks = 260
Inputs / measurements
Wrist-view RGB-D, third-view RGB-D, language, robot action/state
Why this is not directly comparable
This is not the whole training budget: the paper pretrains on LIBERO before Piper fine-tuning. Two RGB-D cameras are documented; no footage was used for this count.
Available artifacts / access
The v3 paper documents the real-robot protocol. This audit did not establish access to the full 260-demonstration corpus.
Source checked
2026-09-29

Isaac GR00T N1.7 training mixture

Official repository / paper
Dataset / version
Isaac GR00T N1.7 training mixture
Training stage
Pretraining mixture
Author-reported scale
20,000 human-video hours reported for N1.7; no combined robot-hour total used here
Inputs / measurements
Images, language, robot state/action, synthetic trajectories
Why this is not directly comparable
This row now identifies N1.7 explicitly. Its human-video quantity is not a total for all GR00T N1 releases and is not robot demonstration time.
Available artifacts / access
N1.7 code and model release are documented. The repository does not establish a single downloadable package for the entire mixed training corpus.
Source checked
2026-09-29

Methodology

These figures answer different questions. LingBot and OpenVLA describe pretraining; TurboVLA's 260 demonstrations describe real-robot fine-tuning after LIBERO pretraining. RoboMIND counts a dataset, not every sample consumed by a model. RDT and OpenVLA draw on overlapping dataset families, so their totals must not be added. Hours are not converted to episodes. A released model or code repository does not establish that its complete training corpus can be downloaded.

Included

Published model, robot, benchmark, deployment or failure evidence that can be tied to a retained source.

Not inferred

Unknown autonomy, latency, trial counts, deployment scale and failure causes remain unknown when the source does not disclose them.

Update policy

Versions and historical records remain distinguishable. A new release does not silently overwrite an older model or evaluation condition.