LATENCY WITHOUT HARDWARE IS MEANINGLESS

Physical AI Inference Latency Leaderboard

Compare published robot-model inference latency with GPU, precision, model size, batch size and test conditions kept visible beside every result.

4 retained recordsLast verified: August 13, 2026Primary-source fields preserved per record
4 of 4 records shownFilters are local and do not create crawlable URLs

TurboVLA

Research result
Model
TurboVLA
Version
216.1M LIBERO configuration
Hardware
NVIDIA RTX 4090
Precision
Paper configuration; precision should be checked in experiment details
Model size
216.1M parameters
Batch
1
Inference latency
31.2 ms per policy inference
Action frequency
~32 policy predictions/s
Control frequency note
Full robot servo loop not claimed to run at 32 Hz
Robot / benchmark
Benchmark + AgileX Piper transfer
Test conditions
Observation-to-action-chunk policy latency on one RTX 4090; excludes full sensing/network/servo pipeline

TurboVLA

Research result
Model
TurboVLA
Version
0.4B RoboTwin bimanual configuration
Hardware
NVIDIA RTX 4090
Precision
Paper configuration
Model size
0.4B parameters
Batch
1 / paper inference setup
Inference latency
43.4 ms
Action frequency
~23 policy predictions/s from latency
Control frequency note
Action chunk = 50 steps; low-level controller frequency is separate
Robot / benchmark
RoboTwin 2.0 bimanual setup
Test conditions
Different model size and observation setup from the 216.1M LIBERO configuration

Isaac GR00T N1

Research paper
Model
Isaac GR00T N1
Version
Public N1 paper model
Hardware
NVIDIA L40
Precision
bf16
Model size
2.2B total; 1.34B VLM reported in paper
Batch
Paper-specific
Inference latency
63.9 ms to sample a chunk of 16 actions
Action frequency
Chunk inference metric; not 16 independent control loops
Control frequency note
Low-level motor control remains separate
Robot / benchmark
Generalist humanoid research model
Test conditions
Published N1 paper, L40 GPU, bf16, 16-action chunk; do not apply directly to N1.7 or other GPUs

WLA-0

Research result
Model
WLA-0
Version
2B active-parameter prototype
Hardware
NVIDIA RTX 5090
Precision
Paper configuration
Model size
2B active parameters
Batch
Paper-specific
Inference latency
40 ms per inference
Action frequency
~25 model inferences/s from reported latency
Control frequency note
Robot low-level control remains separate
Robot / benchmark
Research benchmark platforms
Test conditions
Author-reported on RTX 5090; world prediction can be disabled during inference according to paper

Methodology

Latency rows are not ranked as hardware-independent model scores. GPU/NPU, precision, model version, batch size, action chunk and measurement boundary stay visible so 40 ms on one accelerator is not treated as equivalent to 40 ms on another.

Included

Published model, robot, benchmark, deployment or failure evidence that can be tied to a retained source.

Not inferred

Unknown autonomy, latency, trial counts, deployment scale and failure causes remain unknown when the source does not disclose them.

Update policy

Versions and historical records remain distinguishable. A new release does not silently overwrite an older model or evaluation condition.