| LingBot-VLA 2.0 pretraining mix | Robbyant / Ant Group | China | LingBot-VLA 2.0 | ~60,000 total reported | Not converted from hours | ~50,000 robot trajectory hours | ~10,000 egocentric human video hours | 20 configurations | Vision, language, robot state/action; human egocentric video in pretraining mix | Real robot + egocentric human video | Model/code open; full proprietary pretraining corpus not presented as downloadable dataset | LingBot-VLA 2.0 |
| OpenVLA pretraining data | OpenVLA collaborators | United States | OpenVLA 7B | Not converted from episodes | 970,000 robot episodes reported | Not published as one comparable hour total | Not used as an equivalent robot-hour metric | Multiple | RGB, language and robot actions/state depending on source dataset | Real robot datasets | Open project/model; underlying datasets retain their own access terms | OpenVLA project |
| RDT-1B multi-robot pretraining collection | Tsinghua University / collaborators | China | RDT-1B | Not converted from episodes | 1M+ episodes plus 6K+ ALOHA fine-tuning episodes | Not disclosed | Not disclosed | Multi-robot | Language, up to three RGB views, robot actions/state | Real robot datasets | Code, weights and data resources linked by project | RDT-1B project |
| RoboMIND 2.0 | Beijing Humanoid Robot Innovation Center / collaborators | China | XR-1 ecosystem / MIND-2 research | Not converted from trajectories | 310K+ real-world dual-arm trajectories + 20K simulated trajectories | Not expressed as hours | Not expressed as hours | 6 | Multimodal robot data; includes 12K tactile-enhanced and 20K mobile-manipulation trajectories | Real robot + simulation | Research dataset release | RoboMIND 2.0 paper |
| TurboVLA AgileX Piper task demonstrations | HUST / Huawei | China | TurboVLA | Not disclosed; not estimated from demonstration count | 260 teleoperated demonstrations total across four tasks | Not disclosed | Not disclosed | Single-arm | Wrist-view RGB-D, third-view RGB-D, language, robot action/state | Teleoperation + real robot | Experiment protocol public; raw dataset availability differs from code/checkpoints | TurboVLA project |
| GR00T N1 family training mixture | NVIDIA | United States | Isaac GR00T N1 family | Not disclosed as one comparable aggregate | Version-dependent | Real robot demonstrations included; aggregate not normalized here | Human video included; quantity not normalized here | Multiple | Images, language, robot state/action, synthetic trajectories | Human video + simulation + generated trajectories + real robot | Model/code open for current N1.7 release; full training corpus is not one downloadable dataset | NVIDIA Isaac GR00T |