Blog / Evaluation
Evaluation · 8 min read

Robot evaluation and benchmarking: how do you know a robot is actually getting better?

Loss curves and simulation scores only tell part of the story. Here's how robot policy evaluation works, and what a trustworthy benchmark looks like.

Cindy Dizon

Robot policy evaluation means running a trained robot model through many attempts at a task, scoring each attempt against a written standard, and recording where and why the failures happen. Benchmarking is doing that the same way every time, so one version can be compared with the next. Training curves and simulation scores help, but the only way to know a robot is getting better is to judge what it actually does.

Picture this: your team just trained version 12 of a robot policy. The training curves look great, and the simulation score is up. But does the robot actually fold the towel better than version 3 did?

It’s a harder question than it sounds.

Training a robot is only half the job. The other half is measuring what you built, in a way you can trust and repeat. That’s robot evaluation, and it’s where a lot of progress is either confirmed or quietly misjudged.

A two-armed robot reaching into bins of small boxes on a conveyor line.
One reach into one bin is one rollout. Evaluation is deciding, the same way every time, whether it worked. (Photo: Franck V. / Unsplash)

What Is Robot Policy Evaluation?

In robotics, the trained model that controls the robot is called a policy. One attempt at a task, like picking up a mug, is called a rollout.

Robot policy evaluation means running many rollouts and judging each one. Did the robot succeed? How far did it get? If it failed, where and why? Judging each attempt this way is called robot rollout scoring, and it sits at the center of any serious evaluation.

The idea is the same whether you’re doing VLA model evaluation on a vision-language-action model or embodied AI evaluation on a humanoid. Think of it like grading exams, except the student is a robot and the answers are physical actions.

If you’ve read our post on what makes a robotics dataset usable, you know that training data shapes what a robot can learn. Evaluation is how you find out what it actually learned.

Why Offline Metrics and Simulation Scores Aren’t Enough

It’s tempting to trust the numbers a training run produces. But those numbers don’t always match what happens on a real robot. The SIMPLER study, from researchers at UC San Diego, Stanford, UC Berkeley and Google DeepMind, found that for imitation learning, validation error “is not a good proxy for a policy’s real-world performance.” Across the six policies it ranked, validation error correlated only weakly with real results (an average Pearson score of 0.308, against 0.924 for the study’s own simulated evaluation).

So why not just test on real robots every time? Because real-world robot benchmarking is hard to do well. Rollouts vary from one run to the next, are difficult to reproduce, and are slow to run.

Simulation can help as a complement. PolaRiS, a 2025 project that builds simulated test environments from short video scans of real scenes, reports scores that correlate strongly with RoboArena, a large real-world benchmark. That shows simulation can track reality when it’s built carefully. It doesn’t remove the need to check against real attempts: the authors note that their simulated tests cover only a small subset of what RoboArena tests, and that soft objects and complex contact remain hard to simulate.

Here’s our view: a rising score on a dashboard isn’t evidence that your robot got better. Someone still has to watch it work.

Success Is Rarely Just Pass or Fail

Imagine three rollouts.

In the first, the robot is asked to put a mug on the top shelf. It places the mug upright on the middle shelf. The object is fine, but the task isn’t done. That’s a partial answer.

In the second, the robot grabs a towel corner and the corner slips. It re-grasps, folds the towel, and smooths the edge. The slip is worth logging, but the robot recovered on its own and the fold meets the standard. That’s a pass.

In the third, the robot reaches for a glass on a rack, closes its gripper on nothing, retries from the same pose, and times out. That’s a fail, most likely a missed grasp on a transparent object.

Reasonable people could disagree on the first two. That’s why success rate alone often fails to capture how a policy really performs. A 2025 survey of the reality gap in robotics makes the same point: success metrics are usually binary or aggregate, they don’t reveal where or why something failed, and “two policies with the same success rate might differ significantly in robustness.”

The point isn’t that one call is right. The point is that the same call has to be made every time. Otherwise the number doesn’t mean anything.

What to Measure Beyond Success Rate

Success rate is still the first signal to track, and completion time is a useful second one. Trossen Robotics’ guide to benchmarking on real hardware recommends recording both as primary outcomes while also tracking reliability and safety, which matter more as robots move closer to real-world use.

A good evaluation usually adds:

  • Progress per step: how far the robot got through reach, grasp, lift and place
  • Recovery: whether the robot fixed its own mistake, and how long it took
  • Failure cause: a consistent label for why it failed
  • Where it began: the step where things started to go wrong, not where the robot stopped

Averages can also hide problems. A high mean success rate says little about how a robot behaves in unusual moments, like an unstable grasp or an object in an unexpected position. That’s why corner cases deserve their own attention.

A yellow robot arm mounted on a test bench inside a dark, foam-lined test chamber.
A test only means something if it can be run again. The scene, the objects and the scoring all have to hold still between versions. (Photo: Guille B / Unsplash)

Public Benchmarks vs. Your Own Test Set

Public benchmarks are valuable. Projects like RoboArena, SIMPLER and LIBERO give the field a shared way to compare models.

But they have limits. Benchmarks can saturate, and test scenes can overlap with what models were trained on. NVIDIA’s RoboLab was built partly to address benchmark saturation, visual overlap between training and test scenes, and heavy setup effort.

There’s a bigger issue too. A leaderboard tells you how a model does on someone else’s test. Your robot needs a test built around your deployment: your tasks, your objects, your clutter, your lighting.

A custom test set usually includes scenario catalogues, variations in objects and lighting, and a set of edge cases the policy should face. It also needs a scoring guide written plainly enough that anyone can apply it the same way.

Contact us to learn more about our benchmark and test design.

Find Where Failure Begins, Not Where It Ends

When a robot fails, it’s tempting to record the ending. The robot timed out. The object fell. But the failure usually began earlier, maybe at the grasp, maybe at the first contact.

Robot failure analysis turns a pile of failed videos into something engineers can act on. Each failure is tagged by step, cause and recovery. Over a week, a pattern shows up: wrong target level, corner slips, missed grasps on transparent objects. That tells a team what to fix, not just that something is broken.

Recovery is getting more attention in research, too. RoboRecover, a benchmark published in September 2026, points out that most evaluations focus on the starting scene and the final outcome while paying less attention to what happens during execution. It argues that recovery deserves to be measured as its own dimension of policy evaluation.

How to Run Evaluation That Holds Up Across Versions

Version 12 has to be judged the way version 3 was. Otherwise you can’t tell whether the robot improved or the judging changed.

A few habits make that possible:

  1. A written scoring guide. It defines pass, partial and fail for each task, plus a failure taxonomy. Version it, so changes are tracked.
  2. Calibrated evaluators. Before scoring starts, evaluators work through shared examples until their calls line up.
  3. Double-scoring. A sample of each batch is scored by two people, and the agreement between them is reported. If evaluators disagree often, the guide needs work.
  4. A decision log. Every disputed case and its ruling is recorded, so the same situation gets the same answer next month.
  5. Enough rollouts. A few attempts per task can mislead. You need enough to tell a real difference from luck.
  6. The same standard over time. Every model version gets judged against the same guide.

In real-world testing, success is often judged by people. RoboArena, for example, ranks policies from head-to-head comparisons made by human evaluators at seven institutions. That makes consistency the whole game, and it’s why human evaluation in robotics works best with a dedicated team working from one standard, with the same people scoring month after month.

We covered that trade-off in our post on dedicated data teams vs. crowdsourced labeling.

For Labelix.ai founder Rashid Arif, everyone in robotics talks about making models smarter. Far fewer talk about how you’d know. “If the scoring changes depending on who watched the video, you’re not measuring progress, you’re measuring who was on shift,” he says.

Where Labelix.ai Fits

Labelix.ai provides independent robot evaluation through a dedicated, in-office team: the same trained people on every batch. The team scores every rollout against one written standard, compares policies head to head, and tags failures by cause, step and recovery.

It’s remote-first, so you don’t need to ship hardware. You send recorded rollouts from real robots or simulation, and the team scores them.

Want to know if your robot is really improving? Request a pilot with us today.

A collaborative robot arm placing a small piece on a white table, with rows of finished pieces around it.
Every placement is an attempt someone can score. Enough of them, judged one way, is what turns a demo into a measurement. (Photo: Maria Teneva / Unsplash)
Talk to a human

Have rollouts that need scoring?

Send a representative sample: one task, two policy versions, your hardest episodes. A dedicated Labelix.ai team scores them against a written standard and shows you how the close calls were decided.

FAQs

What is robot policy evaluation?

It's the process of measuring how well a trained robot model performs real tasks. That usually means running many attempts, scoring each one, and analyzing why failures happen.

Why isn't the success rate enough?

Success rate is a useful headline, but it hides partial progress, recovery, and the cause of failures. Two policies can have the same rate and behave very differently.

What is a rollout?

A rollout is one recorded attempt by a robot to complete a task, from start to finish.

How is simulation evaluation different from real-world evaluation?

Simulation is faster, cheaper and easier to repeat, but it can differ from reality. Real-world evaluation is slower and harder to reproduce, but it shows how the robot actually behaves. Many teams use both.

How many trials do you need per task?

There's no single number. You need enough rollouts that a difference between two versions isn't just luck, and tasks with more variation need more. A small pilot is a good way to see how much data your tasks need.

Cindy Dizon
Cindy Dizon
Content Writer · Labelix.ai

Cindy writes about Physical AI, robotics, and the human data that teaches models to perceive and act, for Labelix.ai, an independent data foundry for robotics and multimodal AI.

More from Cindy →
The Data Brief

Get the next one first

One sharp read a month on Physical AI and the data behind it.