BENCHMARK CONDITIONS STAY ATTACHED

Physical AI Benchmark Database

Compare published Physical AI benchmark results while keeping robot, environment, trial count, protocol and adaptation conditions attached to every score.

7 retained recordsLast verified: August 13, 2026Primary-source fields preserved per record
7 of 7 records shownFilters are local and do not create crawlable URLs

TurboVLA 216.1M

Research result
Benchmark
LIBERO (40 tasks across four suites)
Task
General manipulation suites
Model
TurboVLA 216.1M
Robot
LIBERO simulated robot
Environment
Simulation
Trials
2,000 total (50 rollouts × 40 tasks)
Metric
Average success rate
Score
97.7%
Adaptation
Benchmark-trained configuration
Protocol
Author evaluation; same paper protocol for listed baselines

TurboVLA

Research result
Benchmark
AgileX Piper real-robot evaluation
Task
Grab the roller
Model
TurboVLA
Robot
AgileX Piper
Environment
Real lab tabletop
Trials
40
Metric
Success rate
Score
92.5%
Adaptation
Fine-tuned
Protocol
65 teleoperated demonstrations for task; model then evaluated over 40 trials

TurboVLA

Research result
Benchmark
AgileX Piper real-robot evaluation
Task
Move the playing card
Model
TurboVLA
Robot
AgileX Piper
Environment
Real lab tabletop
Trials
40
Metric
Success rate
Score
80%
Adaptation
Fine-tuned
Protocol
65 teleoperated demonstrations for task; model then evaluated over 40 trials

TurboVLA

Research result
Benchmark
AgileX Piper real-robot evaluation
Task
Press the stapler
Model
TurboVLA
Robot
AgileX Piper
Environment
Real lab tabletop
Trials
40
Metric
Success rate
Score
90%
Adaptation
Fine-tuned
Protocol
65 teleoperated demonstrations for task; model then evaluated over 40 trials

TurboVLA

Research result
Benchmark
AgileX Piper real-robot evaluation
Task
Stack three bowls
Model
TurboVLA
Robot
AgileX Piper
Environment
Real lab tabletop
Trials
40
Metric
Success rate
Score
87.5%
Adaptation
Fine-tuned
Protocol
65 teleoperated demonstrations for task; model then evaluated over 40 trials

MotuBrain

Research result
Benchmark
RoboTwin 2.0 Clean
Task
Multi-task robot control
Model
MotuBrain
Robot
RoboTwin 2.0 embodiments
Environment
Simulation benchmark
Trials
See paper protocol
Metric
Average success rate
Score
95.8%
Adaptation
Post-trained
Protocol
Author-reported benchmark; compare only under the same RoboTwin 2.0 setup

WLA-0

Research result
Benchmark
RoboTwin 2.0 Clean
Task
Multi-task / long-horizon learning
Model
WLA-0
Robot
RoboTwin 2.0 embodiments
Environment
Benchmark environment
Trials
See paper protocol
Metric
Success rate
Score
92.94%
Adaptation
Model-specific training
Protocol
Author-reported WLA-0 evaluation

Methodology

Scores are never treated as directly comparable when robots, environments, trial counts, reset rules, benchmark versions or adaptation protocols differ. The table preserves those conditions beside each number.

Included

Published model, robot, benchmark, deployment or failure evidence that can be tied to a retained source.

Not inferred

Unknown autonomy, latency, trial counts, deployment scale and failure causes remain unknown when the source does not disclose them.

Update policy

Versions and historical records remain distinguishable. A new release does not silently overwrite an older model or evaluation condition.