Home/Claim Check

How much of a LIBERO score measures the environment instead of the model

A benchmark is doing two jobs at once: measuring capability and being gamed. When randomisation is thin, the second one wins.

The claim
A high LIBERO score means the model manipulates well, and the leaderboard is a capability ranking.
Reality
A recurring finding in reproduction and evaluation discussion is that a substantial share of samples in this benchmark can be solved by a statistical shortcut: the model never has to learn grasping logic, only the environment textures and where objects usually sit. What it learned is what this environment looks like, not how to pick things up.
Where the gap is
The benchmark is designed to be fair and comparable, but it optimises two objectives at once: measure capability and be gamed. When sampling, textures and initial poses lack sufficient randomisation, the shortcut is the cheapest available solution — not cheating, just an optimiser finding the easier road.
When it closes
With enough randomisation (texture, lighting, initial pose, camera pose) the shortcut stops working and the score returns to the capability itself. Whether a benchmark is trustworthy is decided by how much randomisation it documents, not by where a submission ranks.
This site's call

When you see a LIBERO score, or any simulation score, ask three things: how much initial-pose randomisation, whether texture and lighting are perturbed, and whether there is a test set nobody tuned against. If none can be answered, the number is a pointer, not a basis for selection.

This is not one lab's problem

Benchmarks getting gamed is a general property of evaluation: once the score is the objective, optimisation aims at the score. The differences are only in how well randomisation is done, whether an independent test set exists, and whether anyone asks about conditions when a number is quoted.

For someone building things the practical meaning is simple: a leaderboard decides which model you go and look at; only a real robot decides which one you use. A task that succeeds ten times more often in simulation can collapse to zero on hardware because of one camera calibration.

A more general rule alongside it

Any widely quoted success rate deserves the question "measured under what conditions, over how many trials, and what did the failures look like". That applies equally to papers, vendor demos and crowdfunding pages.

Sources

  • The LIBERO benchmark paper and subsequent reproduction discussion (randomisation settings and shortcut solutions)
  • Reproducibility discussion of vision-language-action models (the distance between simulated scores and real-hardware behaviour)

Objects in this entry

Last checked 2026-09-28