
Introduction
A robot policy's success rate describes a particular set of trials. The scene layouts, object choices, training data and stopping rules all help define that number. A score can rise while the capability a reader cares about remains untested.
A June 2026 study audits five manipulation benchmarks and asks how well their scores support broader claims. It provides a useful way to read Physical AI papers before turning a leaderboard result into a claim about a robot working in a home or factory.
The audit's four diagnostics
The researchers examine shortcut solvability, statistical evidence, repeated adaptation to the test distribution and dependence on the source of training data. They apply these diagnostics to LIBERO, CALVIN, SimplerEnv, RoboCasa and RoboTwin 2.0.
On LIBERO, they report that a 0.09B probe without a language encoder reaches results near leading reported scores. This challenges how much a high aggregate score alone establishes about language-conditioned behavior. It does not show that language is unnecessary for every manipulation task.
Changing the scene inside the training range
The project's CALVIN analysis changes the block poses while remaining within the training range. The reported average number of completed tasks out of five falls from 4.17 to 3.14. That result motivates looking at the distribution of starting states, rather than only the named task suite.
A narrow recurring arrangement can reward a policy that depends on familiar placement. For a reader interested in deployment, the next questions concern the size and direction of position changes, whether objects stay visible and whether the same object identities appear during training. Each condition defines a different generalization claim.
A percentage needs its trial record
Two papers can report the same percentage after using different numbers of attempts, different task mixtures or different reset rules. The percentages alone leave the reader unable to reconstruct the comparison. Publish successes and attempts per task, the tested seeds and the success detector.
For example, 90 successes from 100 attempts and 900 from 1,000 both produce 90 percent. The second sample contains more information about uncertainty under the same sampling assumptions. Repeated runs from one fixed scene also differ from independent runs across varied scenes. More repeats of the same situation do not automatically establish wider coverage.
A comparison between two policies should explain whether they used paired starting states. A failed reset, simulator crash or timed-out episode needs a stated treatment. Excluding difficult episodes after observing the result can alter the reported success rate.
Keep training provenance attached to the score
Record the dataset release, task-specific fine-tuning data, object identities and checkpoint selection procedure. The audit shows why closeness between training examples and test conditions can make a smaller model look competitive with a larger model trained on broader data.
This also affects claims about open models. Releasing weights lets another team run the policy, but a missing training-data description can still prevent that team from interpreting a benchmark gain. Model size, compute and data coverage belong beside the score.
An evaluation card a reader can use
A practical result card would identify the model checkpoint, simulator version, task list, environment seeds, training sources, action frequency and time limit. For each task it would provide successes, attempts, intervention counts and a description of the success detector.
A second evaluation can vary one condition at a time. Move an object, change its identity, alter the camera position or rewrite the instruction while retaining its meaning. Report the size of each change. This proposed reporting format makes failure patterns visible without collapsing every variation into one average.
Simulation remains useful for controlled comparisons. A humanoid working with physical objects also needs tests for contact, sensor delay, actuator limits and recovery. The audit supplies diagnostics for interpreting a score; it does not certify a physical robot for deployment.
Sources and methodology
The findings attributed to the audit come from its authors. The evaluation-card proposal and numerical sampling example are TechniaHQRobot analysis.
Share this article
Share the current TechniaHQRobot article page.
Continue reading
Open the latest robotics reporting, Physical AI analysis and hardware notes.