Is your robot getting better? Someone has to watch.
A dedicated, trained evaluation team that scores every rollout against one written standard, compares policies head to head, and tells your engineers exactly where and why the robot fails.
- Independent & neutral
- One written scoring standard
- Remote-first, no hardware needed
- Never crowdsourced
24 trials, one policy. Each attempt judged against the same written guide.
Offline metrics don't settle it
Loss curves and simulation scores often disagree with what the policy does on a real robot. Someone still has to judge real attempts.
Success is rarely binary
Wrong shelf, tipped object, recovered slip, task done out of order. Each needs the same call every time, or the number means nothing.
Every version, same standard
Version 12 has to be judged the way version 3 was. That takes a written guide and the same evaluators, not whoever is free this week.
Score three rollouts. Then compare with the guide.
Each rollout is a real kind of attempt a household robot makes. Make your call, then see how a written scoring guide decides it, and what gets tagged. Most teams disagree on at least one.
Put the mug on the top shelf
- 0:01reach
- 0:03grasp mug
- 0:06lift
- 0:09place on middle shelf
- 0:10release
Fold the towel in half
- 0:02grasp corner
- 0:04corner slips
- 0:05re-grasp
- 0:08fold
- 0:10smooth edge
Pick the glass from the rack
- 0:01reach
- 0:03close gripper
- 0:04no contact
- 0:07retry, same pose
- 0:60timeout
Score all three to compare with the guide.
A score is the headline. The tags are the story.
Reach, grasp, lift, place. Where each phase of the attempt starts and ends.
Contact, slip, re-grasp, drop, each marked at the frame it happened.
Not where the robot gave up. Where the failure actually started.
Why it failed, from your agreed taxonomy, and whether the robot recovered on its own.
Three ways in. One team that learns your robot.
Rollout scoring & policy comparison
Every attempt scored pass, partial or fail against your guide, with progress per step. Two policies compared side by side on the same tasks.
- Success, partial, failure
- Progress per step
- Head-to-head comparisons
- Release-gate reports
Failure analysis
Every failed and borderline attempt tagged by cause, step and recovery, rolled up into a weekly breakdown of where the policy breaks.
- Failure cause taxonomy
- Where failure began
- Recovery & interventions
- Weekly breakdowns
Benchmark & test design
The scenarios, variations and edge cases your policy should face, and a scoring guide written plainly enough to apply the same way every time.
- Scenario catalogues
- Object, clutter, lighting variation
- Edge-case suites
- Versioned scoring guides
A report your engineers can act on on Monday.
Outcome mix per task, the failure cause that matters most, and how often our evaluators agreed with each other. Per model version, so you can see what actually changed.
| Task | Outcome mix | Top failure cause |
|---|---|---|
| Mug to shelf | wrong target level | |
| Fold towel | corner slip | |
| Glass from rack | missed grasp, transparent | |
| Open drawer | handle collision |
Two weeks from rollouts to a decision.
You send rollouts
Recorded attempts from one policy, real or sim, plus your current pass criteria if you have them.
We draft the guide
A written scoring guide and failure taxonomy, built from your footage and agreed with your engineers.
Calibrate, then score
Evaluators calibrate on shared examples, then score every rollout. A sample of every batch is scored twice.
Report and agreement
Scores, failure breakdown, and our evaluator agreement numbers. You decide whether to scale.
Remote evaluation
Recorded rollouts from real robots or simulation, scored, compared and tagged by our team. No hardware needed on our side.
Live trials
Supervising and resetting your robots over a remote link, or a test bench staged for your tasks in our Dhaka facility. Planned together, not quoted off a menu.
Questions, answered straight.
Robot policy evaluation measures how well a trained robot model performs real tasks. It usually means running many attempts per task and judging each one: did it succeed, how far did it get, and if it failed, why. Automated metrics help, but the final judgment on real attempts is still made by people, which is the work Labelix.ai staffs.
Offline metrics and simulation scores often correlate poorly with how a policy behaves on a real robot, and many outcomes are not binary. A partial fold, a recovered grasp or a task done in the wrong order needs a consistent human call. We apply one written standard so those calls stay comparable across model versions.
Not to start. Most teams begin with remote evaluation of recorded rollouts, from real robots or simulation, which needs no hardware on our side. Live trial support, on your robots over a remote link or on a test bench staged for your project, is scoped with you when you need it.
A written scoring guide agreed with your team, calibration on shared examples before scoring starts, double-scoring on a sample of every batch, and a decision log for every disputed case. You see the agreement numbers, not just the scores.
With a small paid pilot: one policy, a set of recorded rollouts and a scoring guide we draft together. Pricing depends on volume and how detailed the scoring is, so we scope it with you rather than quote a fixed number. The pilot shows you the output and our agreement numbers before you scale.
Send us a week of rollouts. Get back a scored, explained report.
Start with a small paid pilot on recorded trials from one policy. You get every rollout scored against a standard you approve, a failure breakdown, and the agreement numbers between our evaluators, before you commit to anything larger.