Datasets & Models
LIBERO benchmark
One of the most-used simulation benchmarks for vision-language-action models. Its score curves look great, which is exactly why it is where the argument about 'what the score actually measures' concentrates: it is doing two jobs at once — measuring capability and being gamed.