WHAT ACTUALLY RAN ON A REAL ROBOT

Real-Robot Evidence Database

See what Physical AI models actually demonstrated on real robots, including task, control mode, trials, failures, interventions and evidence limits.

5 retained recordsLast verified: August 13, 2026Primary-source fields preserved per record
5 of 5 records shownFilters are local and do not create crawlable URLs

TurboVLA

Research result
Model
TurboVLA
Robot
AgileX Piper
Task
Four tabletop tasks: roller, playing card, stapler, bowl stacking
Environment
Lab tabletop
Control mode
Autonomous policy execution after teleoperated training data collection
Autonomy level
Task policy
Trials
160 total
Success
87.5% aggregate across the four 40-trial task sets
Failures
20 total inferred exactly from published trial counts and success rates
Human interventions
Runtime intervention count not reported
Limitations
Controlled tabletop tasks on one arm; no long-duration deployment evidence

XR-1

Research result
Model
XR-1
Robot
Six robot embodiments
Task
>120 manipulation tasks
Environment
Real-world research evaluations
Control mode
Learned VLA policy
Autonomy level
Autonomous evaluation under paper protocol
Trials
>14,000 rollouts reported
Success
Not reduced to one cross-task figure here
Failures
Model/task-specific; no single aggregate count represented here
Human interventions
See paper protocol
Limitations
Published research evaluation, not proof of unsupervised production deployment

RDT-1B

Research result
Model
RDT-1B
Robot
ALOHA dual-arm and other manipulation platforms
Task
Bimanual manipulation, language following, unseen objects/scenes and few-shot skills
Environment
Real robot research evaluation
Control mode
Diffusion action policy
Autonomy level
Autonomous policy evaluation under paper protocol
Trials
See paper
Success
Not collapsed into one universal number
Failures
See task-specific tables
Human interventions
Not normalized here
Limitations
Research task protocols differ from deployment workloads

Gemini Robotics 2

Official demo
Model
Gemini Robotics 2
Robot
Apptronik Apollo
Task
Whole-body locomotion plus manipulation demonstrations
Environment
Real robot demonstrations
Control mode
VLA whole-body control
Autonomy level
Official autonomous task demonstrations; exact supervision varies by experiment
Trials
Not disclosed as one common total
Success
Task-specific official results
Failures
Not disclosed as one common total
Human interventions
Not disclosed as one common rate
Limitations
Closed model and controlled evaluation; public data does not establish fleet reliability

OpenVLA 7B

Research result
Model
OpenVLA 7B
Robot
Multiple manipulation robots
Task
Language-conditioned manipulation
Environment
Real robot research evaluations
Control mode
VLA policy
Autonomy level
Autonomous policy evaluation under project protocol
Trials
Protocol-dependent
Success
See individual project evaluations
Failures
Protocol-dependent
Human interventions
Not normalized across platforms
Limitations
Different robots require adaptation; benchmark and real-world protocols are not interchangeable

Methodology

Scripted demos, teleoperation, shared autonomy, autonomous policy tests and deployments are kept separate. Training teleoperation is not mislabeled as runtime teleoperation, and no system is called fully autonomous without explicit evidence.

Included

Published model, robot, benchmark, deployment or failure evidence that can be tied to a retained source.

Not inferred

Unknown autonomy, latency, trial counts, deployment scale and failure causes remain unknown when the source does not disclose them.

Update policy

Versions and historical records remain distinguishable. A new release does not silently overwrite an older model or evaluation condition.