Training data

Robot Learning Data Pipeline

Fourteen official volume records now cover teleoperation hours, human-motion data, trajectories, collection robots, tasks and simulation environments. The chart groups values by unit so one million trajectories are not treated as one million hours.

Data summary

What is included

Map the stages between teleoperation, sensor capture, cleaning, annotation, training, validation and robot deployment.

Visible rows

14

Dataset scope

TechniaHQRobot source-linked dataset

Interactive infographic

Robot Learning Data Pipeline

Pipeline Sankey · Validation funnel · Collected vs retained bars. Exact values, evidence status and URLs are repeated in the accessible table.

Accessible data table

The table contains the same source-linked records used by the chart. It is the reference when labels overlap or a visual scale compresses values.

Source-linked data used in the infographic
Pipeline stageOrderDocumented volumeUnitStatusEvidenceSource
Figure Helix multi-robot teleoperation dataset1500hoursofficially-documentedOfficial announcementFigure Helix official technical articleChecked 2026-07-31
Figure Helix logistics training subset210hoursofficially-documentedOfficial announcementFigure official Helix logistics scaling studyChecked 2026-07-31
Figure Helix logistics training subset220hoursofficially-documentedOfficial announcementFigure official Helix logistics scaling studyChecked 2026-07-31
Figure Helix logistics training subset240hoursofficially-documentedOfficial announcementFigure official Helix logistics scaling studyChecked 2026-07-31
Figure Helix logistics training subset260hoursofficially-documentedOfficial announcementFigure official Helix logistics scaling studyChecked 2026-07-31
Figure Helix 02 human-motion retargeting31000hoursofficially-documentedOfficial announcementFigure Helix 02 official technical articleChecked 2026-07-31
Figure Helix 02 parallel simulation4200000parallel environmentsofficially-documentedOfficial announcementFigure Helix 02 official technical articleChecked 2026-07-31
AgiBot World embodied dataset51000000trajectoriesofficially-documentedOfficial announcementAgiBot official research pageChecked 2026-07-31
AgiBot World collection fleet5100robotsofficially-documentedOfficial announcementAgiBot official research pageChecked 2026-07-31
AgiBot World task coverage51000tasksofficially-documentedOfficial announcementAgiBot official research pageChecked 2026-07-31
Fourier ActionNet bimanual teleoperation dataset630000trajectoriesofficially-documentedOfficial announcementFourier official teleoperation and ActionNet documentationChecked 2026-07-31
1X egocentric human video7900hoursofficially-documentedOfficial announcement1X official world-model training articleChecked 2026-07-31
1X robot fine-tuning data870hoursofficially-documentedOfficial announcement1X official world-model training articleChecked 2026-07-31
1X unfiltered robot data8400hoursofficially-documentedOfficial announcement1X official world-model training articleChecked 2026-07-31

Methodology

How the dataset is built

Hours, trajectories, robots, tasks and simulation environments are different units and are never summed into one volume. Each row preserves the collection modality and scope stated by the source.

Editorial analysis

Scope and direct answer

This page is built around a single measurable question rather than a broad ranking. The infographic, accessible table and downloadable CSV use the same local rows, so the visual cannot quietly introduce a value that is absent from the dataset. The direct answer above states the current scope, while the evidence badges show how each row was documented. A missing field remains null and is displayed as unavailable instead of being converted into zero, an average or a guessed category.

The dataset is intentionally narrower than the humanoid robotics industry. It describes only entries that can be tied to a source URL and interpreted under the page’s definitions. That design makes the page slower to fill but easier to audit. Readers can open the source, check the date, inspect the measurement condition and decide whether two rows are genuinely comparable before drawing a conclusion.

Dataset boundaries and inclusion rules

A row is included when the robot, company, event or metric can be identified and the source supplies enough context to preserve its meaning. Manufacturer product pages, technical documentation, manuals, original research papers, regulatory filings, official company announcements and customer deployment pages receive priority. Secondary reporting can add context, but it should not replace a primary source when the primary document is available.

The dataset excludes unattributed screenshots, recycled specification tables, numerical claims without a date and values derived from appearance alone. A promotional video can document that a motion or task occurred under the shown conditions, but it cannot establish hidden payload, autonomy, production volume or long-term reliability. When evidence is incomplete, the row can remain in draft with an explicit note rather than being upgraded to a stronger status.

Definitions and comparable units

Comparable charts depend on definitions that remain stable from one row to the next. Each field has a specific label, unit and condition. The local schema used here includes: Pipeline stage, Order, Documented volume, Unit, Status. Values with the same label can still describe different tests, so notes and measurement conditions are treated as part of the metric rather than optional commentary.

Units are never stripped to simplify the design. Original currencies remain attached to prices, time units remain attached to runtime or session data, and force or mass values retain their stated convention. Conversions can be displayed temporarily when the method and date are visible, but the source value stays unchanged in the JSON file. This prevents later updates from compounding an old conversion or mixing incompatible systems.

How to read the visual

The requested visual forms for this topic are Pipeline Sankey, Validation funnel, Collected vs retained bars. The component chooses the clearest form supported by the rows currently loaded. A categorical bar chart emphasizes counts or source values; a scatter plot exposes relationships between two numeric fields; a timeline preserves event order; a matrix makes presence and absence visible; an architectural diagram shows sequence or structure without pretending that a percentage exists.

Filters change only the visible subset. They do not rewrite source values, recalculate missing entries or merge categories behind the reader’s back. Exact values remain available in the HTML table, which is also the safer reference on small screens or when labels overlap. The chart uses labels and shapes in addition to color so that meaning does not depend on a single visual channel.

Evidence hierarchy and source handling

Every row carries an evidence status such as official specification, official announcement, customer confirmed, regulatory document, scientific paper, third-party report, video observation or not disclosed. The badge does not declare that a source is infallible. It tells the reader what kind of claim is being made and how directly the source supports it. A manufacturer can accurately publish a specification while leaving the test protocol undocumented.

Source dates and access dates serve different purposes. The source date shows when the claim was published or measured; the access date records when TechniaHQRobot last checked it. Product pages can change without preserving earlier versions, so the notes field should capture configuration, wording or conditions that could affect future verification. When a source disappears, the row should not silently inherit a new URL with different content.

Technical interpretation

Teleoperation, sensor capture, synchronization, cleaning, segmentation, annotation, filtering, training, simulation, robot validation, deployment and field feedback form distinct pipeline stages. Recorded hours, retained trajectories, training episodes and validated tasks are different units. A large raw collection can shrink after quality control without implying that the discarded data was useless.

Volumes should appear only when a primary source or another documented source reports them. The architecture can remain qualitative when companies do not disclose internal data quantities. These distinctions determine which rows may share one scale and which need separate filters, symbols or tables. The purpose of the infographic is not to force every robot into a universal score. It is to make the structure of the evidence visible enough that a reader can identify where a comparison is strong, weak or impossible.

Bias, missing data and uncertainty

Public humanoid robotics data is selective. Companies tend to publish the specifications that support a product narrative, while unsuccessful trials, integration labor, intervention frequency and maintenance history are less likely to appear. The resulting dataset can overrepresent commercially promoted platforms and underrepresent university systems, private pilots or discontinued programs. A larger number of rows therefore does not automatically mean a more complete view of the market.

Missingness can itself be informative, but it must not be turned into a hidden score. A null value can mean the manufacturer did not publish the metric, the source uses an incompatible definition, the product is still a prototype or verification is pending. The page displays those states directly. It does not penalize an entry numerically or assume that an undisclosed capability is absent.

What the figures cannot prove

A chart can organize documented values; it cannot prove general autonomy, safe unsupervised operation, mass production, commercial success or reliable performance in every environment. A high specification does not guarantee good control software. A large funding round does not prove hardware readiness. A customer pilot does not reveal the number of interventions unless the customer or manufacturer publishes that evidence.

Patterns should be treated as questions for further investigation. A correlation can arise from product category, configuration, release date, source selection or missing data. The analysis should remain proportional to the sample and avoid language that converts a limited source-linked dataset into a global industry statistic. Where the denominator is the TechniaHQRobot catalog, that denominator is stated beside the chart.

Maintaining the dataset

New records should preserve the existing fields and include the robot or company name, original value, unit, measurement condition, source name, absolute source URL, source date, access date, evidence status and a concise note when context is required. Unavailable values should remain empty rather than being replaced with zero.

Publication checks

Before publication, confirm that every numeric value has a source, all units are visible, and the chart uses only fields that share a defensible definition. Test filters with keyboard navigation, review the table with a screen reader, inspect the layout on a narrow mobile viewport and print the page to verify that titles, units, methodology and sources remain attached to the visual. Check that the canonical URL matches the preserved route.

The comparison should remain limited when the sample is too sparse to support the promised analysis. Adding one number is not enough: the dataset needs a clear denominator, compatible conditions and enough source coverage to avoid a misleading visual. Update dates should change only when the content or data actually changes, and internal links should connect the page to the main humanoid guide, catalog and relevant robot profiles.

Technical fields used by this page

The dataset includes Pipeline stage, Order, Documented volume, Unit, Status. A field can remain unavailable when the manufacturer does not publish it or when the available source measures a different condition. This prevents a visually complete chart from becoming a technically false comparison.

The requested visual forms are Pipeline Sankey, Validation funnel, Collected vs retained bars. The current component selects the clearest representation supported by the loaded rows. When the dataset lacks a defensible numeric series, it displays an architecture or empty state instead of manufacturing percentages.

Frequently asked questions

What does the robot learning data pipeline page measure?

Map the stages between teleoperation, sensor capture, cleaning, annotation, training, validation and robot deployment.

How are missing values handled?

Missing values are stored as null and displayed as not publicly disclosed, not specified by the manufacturer, no comparable public data or awaiting verified data. They are never replaced with zero.

Can the local chart be treated as a global market statistic?

No. The chart describes only the source-linked rows in the TechniaHQRobot source-linked dataset unless the page explicitly cites a broader primary dataset.

Sources

Related humanoid infographics

Continue with the verified humanoid catalog

Open individual robot records for manufacturer, origin, dimensions, availability, public price status and official product links.