Physical AI data
Reading time 7 min readManiformer

Maniformer MEgo reaches 20,000 devices and one million hours of Physical AI data

Maniformer says 20,000 MEgo devices have produced more than one million hours of real-world data. We examine coverage, transfer value and quality checks.

By TechniaHQRobot

Introduction

Maniformer, the embodied AI data company incubated by AgiBot, says its 20,000th MEgo data-collection device rolled off the line on August 31 and was delivered to JD.com. The company also reports more than one million hours of high-quality embodiment-free data collected in open real environments.

The scale deserves a closer look because robot-learning data is only useful when its coverage, labels, sensor quality and transfer path are known. One million hours can contain broad behavioral diversity, repeated low-value actions or a mixture of both. The useful question for a model developer is how much of that data transfers into better policy performance on a target robot.

Information verified from official sources available as of September 24, 2026.

Checks that make a large Physical AI dataset easier to evaluate

AreaUseful metricReason
CoverageHours by task, scene and object familyShows whether a large total is concentrated in a few common tasks
Sensor integrityDropped-frame and synchronization-error rateShows how much recorded data is technically usable
Operator diversityNumber of operators and variation in action styleHelps estimate behavior diversity
TransferRobot success after adding MEgo dataConnects human demonstrations to robot performance
EfficiencyTarget-robot hours needed after pretrainingTests whether external data reduces expensive robot collection

These are TechniaHQRobot evaluation criteria. Maniformer has not published all of these metrics.

The dataset spans many environments and task types

Chinese reporting based on Maniformer's announcement says the corpus covers 22 scene categories, more than 10,000 real environments, over 50,000 object types and more than 500 fine-grained tasks. The company says the data comes from open real environments rather than simulation alone.

Maniformer's website describes three collection paths: real robot data, simulation data and human demonstration data. MEgo belongs to the human-centered collection side, using wearable or handheld devices to record interactions before mapping them into robot training pipelines.

Those counts describe breadth. They do not reveal how evenly the hours are distributed across tasks, how many environments appear only once, or how much of the corpus contains rare recovery behavior. A training team needs that distribution before deciding whether raw hours represent useful coverage.

Embodiment-free data still needs a transfer layer

MEgo Gripper is designed to record human grasping, placement and manipulation with cameras and motion sensors. The advantage is collection speed without occupying a full robot for every demonstration.

Human motion does not map directly to every robot hand or arm. Joint limits, gripper geometry, camera placement, reach and force capability differ across platforms. A training pipeline still needs retargeting or a model that can learn an action representation shared across embodiments.

A useful evaluation compares two policies trained with the same amount of target-robot data, then adds MEgo data to one of them. Report success rate, required robot demonstrations and performance after object position, camera angle or task order changes.

One million hours should be accompanied by data-quality statistics

For large Physical AI datasets, recording failures in the data pipeline matters as much as counting captured hours. Useful statistics include dropped frames, camera synchronization error, missing sensor channels, failed uploads and segments rejected during quality control.

Task metadata also matters. A clip labeled 'pick object' can involve different approach directions, grip widths, contact forces and destinations. Training value improves when the model can distinguish those conditions.

Maniformer's scale gives researchers a reason to watch this dataset. Public benchmark results showing how the corpus changes downstream robot success would make the scale easier to compare with other data providers.

The Tencent and JD links point to two different uses

Maniformer announced a data-business collaboration with Tencent Robotics X when the 20,000th MEgo unit was delivered to JD.com. One relationship concerns model and research data. The other shows a large commercial customer receiving collection hardware.

For buyers, the practical question is whether data collection stays consistent across many operators and sites. Calibration drift, different handling styles and incomplete task instructions can add variation that a model may or may not use productively.

A strong multi-site report would publish inter-device calibration checks, rejected-data rates and downstream transfer results by task family. That would connect collection scale to model behavior.

Limitations and missing information

  • The one-million-hour figure and environment counts are company-reported and were not independently audited in the sources reviewed.
  • Embodiment-free data volume does not by itself establish better downstream robot performance. Transfer quality needs task-level testing.

Sources and methodology

Share this article

Share the current TechniaHQRobot article page.

Continue reading

Open the latest robotics reporting, Physical AI analysis and hardware notes.

Browse robotics news
Article by @techniahqrobot