| LingBot-VLA 2.0 pretraining mix | Pretraining mixture | ~50,000 robot hours + ~10,000 egocentric human-video hours | Vision, language, robot state/action; human egocentric video in pretraining mix | Human video and robot-action data are distinct components. The 60,000-hour total does not mean 60,000 robot-control hours. | Repository links code and model artifacts. It does not present the complete proprietary pretraining corpus as a single downloadable dataset. | 2026-09-29 | LingBot-VLA 2.0 |
| OpenVLA pretraining data | Pretraining mixture | 970,000 robot episodes from Open X-Embodiment | RGB, language and robot actions/state depending on source dataset | Episode duration varies and no comparable hour total is given. Some underlying datasets also appear in other mixtures. | Model/code are released; underlying Open X-Embodiment datasets have separate access and licensing terms. | 2026-09-29 | OpenVLA project |
| RDT-1B multi-robot pretraining collection | Pretraining and fine-tuning | 1M+ pretraining episodes from 46 datasets; separately, 6K+ ALOHA fine-tuning episodes | Language, up to three RGB views, robot actions/state | The 46-dataset mix and the authors' ALOHA collection describe separate stages. Shared dataset families prevent adding this budget to OpenVLA's. | Code, weights and the authors' ALOHA fine-tuning dataset are linked. The model card's MIT statement is not a blanket licence for all 46 source datasets. | 2026-09-29 | RDT-1B projectRDT-1B model cardALOHA fine-tuning dataset |
| RoboMIND 2.0 | Dataset release | 310K+ real-world trajectories; an additional 20K simulated trajectories | Multimodal robot data; includes 12K tactile-enhanced and 20K mobile-manipulation trajectories | The 12K tactile and 20K mobile-manipulation trajectories are subsets of the real-world data, not extra totals. Six embodiments and 739 tasks describe coverage. | The paper describes the dataset. Complete file availability and dataset licence were not established by this audit. | 2026-09-29 | RoboMIND 2.0 paper, v3 |
| TurboVLA AgileX Piper task demonstrations | Real-robot fine-tuning | 65 teleoperated demonstrations per task × 4 tasks = 260 | Wrist-view RGB-D, third-view RGB-D, language, robot action/state | This is not the whole training budget: the paper pretrains on LIBERO before Piper fine-tuning. Two RGB-D cameras are documented; no footage was used for this count. | The v3 paper documents the real-robot protocol. This audit did not establish access to the full 260-demonstration corpus. | 2026-09-29 | TurboVLA paper, v3 (27 September 2026)TurboVLA project (different results version) |
| Isaac GR00T N1.7 training mixture | Pretraining mixture | 20,000 human-video hours reported for N1.7; no combined robot-hour total used here | Images, language, robot state/action, synthetic trajectories | This row now identifies N1.7 explicitly. Its human-video quantity is not a total for all GR00T N1 releases and is not robot demonstration time. | N1.7 code and model release are documented. The repository does not establish a single downloadable package for the entire mixed training corpus. | 2026-09-29 | NVIDIA Isaac GR00T |