Introduction
Maniformer, the embodied AI data company incubated by AgiBot, says its 20,000th MEgo data-collection device rolled off the line on August 31 and was delivered to JD.com. The company also reports more than one million hours of high-quality embodiment-free data collected in open real environments.
The scale deserves a closer look because robot-learning data is only useful when its coverage, labels, sensor quality and transfer path are known. One million hours can contain broad behavioral diversity, repeated low-value actions or a mixture of both. The useful question for a model developer is how much of that data transfers into better policy performance on a target robot.
Information verified from official sources available as of September 24, 2026.
Checks that make a large Physical AI dataset easier to evaluate
| Area | Useful metric | Reason |
|---|---|---|
| Coverage | Hours by task, scene and object family | Shows whether a large total is concentrated in a few common tasks |
| Sensor integrity | Dropped-frame and synchronization-error rate | Shows how much recorded data is technically usable |
| Operator diversity | Number of operators and variation in action style | Helps estimate behavior diversity |
| Transfer | Robot success after adding MEgo data | Connects human demonstrations to robot performance |
| Efficiency | Target-robot hours needed after pretraining | Tests whether external data reduces expensive robot collection |
These are TechniaHQRobot evaluation criteria. Maniformer has not published all of these metrics.
The dataset spans many environments and task types
Chinese reporting based on Maniformer's announcement says the corpus covers 22 scene categories, more than 10,000 real environments, over 50,000 object types and more than 500 fine-grained tasks. The company says the data comes from open real environments rather than simulation alone.
Maniformer's website describes three collection paths: real robot data, simulation data and human demonstration data. MEgo belongs to the human-centered collection side, using wearable or handheld devices to record interactions before mapping them into robot training pipelines.
Those counts describe breadth. They do not reveal how evenly the hours are distributed across tasks, how many environments appear only once, or how much of the corpus contains rare recovery behavior. A training team needs that distribution before deciding whether raw hours represent useful coverage.
Embodiment-free data still needs a transfer layer
MEgo Gripper is designed to record human grasping, placement and manipulation with cameras and motion sensors. The advantage is collection speed without occupying a full robot for every demonstration.
Human motion does not map directly to every robot hand or arm. Joint limits, gripper geometry, camera placement, reach and force capability differ across platforms. A training pipeline still needs retargeting or a model that can learn an action representation shared across embodiments.
A useful evaluation compares two policies trained with the same amount of target-robot data, then adds MEgo data to one of them. Report success rate, required robot demonstrations and performance after object position, camera angle or task order changes.
One million hours should be accompanied by data-quality statistics
For large Physical AI datasets, recording failures in the data pipeline matters as much as counting captured hours. Useful statistics include dropped frames, camera synchronization error, missing sensor channels, failed uploads and segments rejected during quality control.
Task metadata also matters. A clip labeled 'pick object' can involve different approach directions, grip widths, contact forces and destinations. Training value improves when the model can distinguish those conditions.
Maniformer's scale gives researchers a reason to watch this dataset. Public benchmark results showing how the corpus changes downstream robot success would make the scale easier to compare with other data providers.
The Tencent and JD links point to two different uses
Maniformer announced a data-business collaboration with Tencent Robotics X when the 20,000th MEgo unit was delivered to JD.com. One relationship concerns model and research data. The other shows a large commercial customer receiving collection hardware.
For buyers, the practical question is whether data collection stays consistent across many operators and sites. Calibration drift, different handling styles and incomplete task instructions can add variation that a model may or may not use productively.
A strong multi-site report would publish inter-device calibration checks, rejected-data rates and downstream transfer results by task family. That would connect collection scale to model behavior.
Limitations and missing information
- The one-million-hour figure and environment counts are company-reported and were not independently audited in the sources reviewed.
- Embodiment-free data volume does not by itself establish better downstream robot performance. Transfer quality needs task-level testing.
Sources and methodology
Share this article
Share the current TechniaHQRobot article page.
Continue reading
Open the latest robotics reporting, Physical AI analysis and hardware notes.