Robot evaluation & benchmarking

Is your robot getting better? Someone has to watch.

A dedicated, trained evaluation team that scores every rollout against one written standard, compares policies head to head, and tells your engineers exactly where and why the robot fails.

Score a rollout yourself ↓
  • Independent & neutral
  • One written scoring standard
  • Remote-first, no hardware needed
  • Never crowdsourced

24 trials, one policy. Each attempt judged against the same written guide.

01

Offline metrics don't settle it

Loss curves and simulation scores often disagree with what the policy does on a real robot. Someone still has to judge real attempts.

02

Success is rarely binary

Wrong shelf, tipped object, recovered slip, task done out of order. Each needs the same call every time, or the number means nothing.

03

Every version, same standard

Version 12 has to be judged the way version 3 was. That takes a written guide and the same evaluators, not whoever is free this week.

Try it

Score three rollouts. Then compare with the guide.

Each rollout is a real kind of attempt a household robot makes. Make your call, then see how a written scoring guide decides it, and what gets tagged. Most teams disagree on at least one.

Rollout Atask

Put the mug on the top shelf

  1. 0:01reach
  2. 0:03grasp mug
  3. 0:06lift
  4. 0:09place on middle shelf
  5. 0:10release
Your call
Rollout Btask

Fold the towel in half

  1. 0:02grasp corner
  2. 0:04corner slips
  3. 0:05re-grasp
  4. 0:08fold
  5. 0:10smooth edge
Your call
Rollout Ctask

Pick the glass from the rack

  1. 0:01reach
  2. 0:03close gripper
  3. 0:04no contact
  4. 0:07retry, same pose
  5. 0:60timeout
Your call

Score all three to compare with the guide.

Inside one scored rollout

A score is the headline. The tags are the story.

01
Step segments

Reach, grasp, lift, place. Where each phase of the attempt starts and ends.

02
Events

Contact, slip, re-grasp, drop, each marked at the frame it happened.

03
Where it began

Not where the robot gave up. Where the failure actually started.

04
Cause & recovery

Why it failed, from your agreed taxonomy, and whether the robot recovered on its own.

What we do

Three ways in. One team that learns your robot.

Rollout scoring & policy comparison

Every attempt scored pass, partial or fail against your guide, with progress per step. Two policies compared side by side on the same tasks.

  • Success, partial, failure
  • Progress per step
  • Head-to-head comparisons
  • Release-gate reports
Start with a pilot →

Failure analysis

Every failed and borderline attempt tagged by cause, step and recovery, rolled up into a weekly breakdown of where the policy breaks.

  • Failure cause taxonomy
  • Where failure began
  • Recovery & interventions
  • Weekly breakdowns
Failure analysis →

Benchmark & test design

The scenarios, variations and edge cases your policy should face, and a scoring guide written plainly enough to apply the same way every time.

  • Scenario catalogues
  • Object, clutter, lighting variation
  • Edge-case suites
  • Versioned scoring guides
Test design →
What you get back

A report your engineers can act on on Monday.

Outcome mix per task, the failure cause that matters most, and how often our evaluators agreed with each other. Per model version, so you can see what actually changed.

Policy v12 · weekly evalexample layout, not client data
TaskOutcome mixTop failure cause
Mug to shelf
wrong target level
Fold towel
corner slip
Glass from rack
missed grasp, transparent
Open drawer
handle collision
passpartialfail+ evaluator agreement on a double-scored sample
The pilot

Two weeks from rollouts to a decision.

1
Week 0

You send rollouts

Recorded attempts from one policy, real or sim, plus your current pass criteria if you have them.

2
Week 1

We draft the guide

A written scoring guide and failure taxonomy, built from your footage and agreed with your engineers.

3
Week 1-2

Calibrate, then score

Evaluators calibrate on shared examples, then score every rollout. A sample of every batch is scored twice.

4
Week 2

Report and agreement

Scores, failure breakdown, and our evaluator agreement numbers. You decide whether to scale.

Available now

Remote evaluation

Recorded rollouts from real robots or simulation, scored, compared and tagged by our team. No hardware needed on our side.

Scoped with you

Live trials

Supervising and resetting your robots over a remote link, or a test bench staged for your tasks in our Dhaka facility. Planned together, not quoted off a menu.

FAQ

Questions, answered straight.

Robot policy evaluation measures how well a trained robot model performs real tasks. It usually means running many attempts per task and judging each one: did it succeed, how far did it get, and if it failed, why. Automated metrics help, but the final judgment on real attempts is still made by people, which is the work Labelix.ai staffs.

Offline metrics and simulation scores often correlate poorly with how a policy behaves on a real robot, and many outcomes are not binary. A partial fold, a recovered grasp or a task done in the wrong order needs a consistent human call. We apply one written standard so those calls stay comparable across model versions.

Not to start. Most teams begin with remote evaluation of recorded rollouts, from real robots or simulation, which needs no hardware on our side. Live trial support, on your robots over a remote link or on a test bench staged for your project, is scoped with you when you need it.

A written scoring guide agreed with your team, calibration on shared examples before scoring starts, double-scoring on a sample of every batch, and a decision log for every disputed case. You see the agreement numbers, not just the scores.

With a small paid pilot: one policy, a set of recorded rollouts and a scoring guide we draft together. Pricing depends on volume and how detailed the scoring is, so we scope it with you rather than quote a fixed number. The pilot shows you the output and our agreement numbers before you scale.

Send us a week of rollouts. Get back a scored, explained report.

Start with a small paid pilot on recorded trials from one policy. You get every rollout scored against a standard you approve, a failure breakdown, and the agreement numbers between our evaluators, before you commit to anything larger.