Robot benchmark & test design

Test the robot you'll ship, not the demo you filmed.

We design the scenarios, variations and edge cases your policy should face, write a scoring guide plain enough to apply the same way every time, and keep it versioned as your robot and your product change.

  • Independent & neutral
  • Written, versioned protocols
  • Built with your engineers
  • Never crowdsourced

One base task fanned out across object, clutter and lighting, with edge cases marked and the scoring guide written.

01

Demo setups flatter policies

A policy tuned on three favourite scenes can look strong there and fall apart on the first unfamiliar kitchen.

02

Variation is a design choice

Every object, clutter level and light condition multiplies the suite. Picking the combinations that matter is the work.

03

Rules in someone's head don't rerun

If the pass criteria aren't written down with examples, results change when the person running the test changes.

Try it

Build a test suite. Watch it grow.

Pick the variations for one household task. See how fast the suite multiplies, which combinations become edge cases, and how many trials it takes to score it.

Base task
Put the item on the shelf
Object
Clutter
Lighting
Start pose
8
scenarios
2
edge cases
80
trials to score
mug · clear
day · upright
mug · clear
dim · upright
mug · heavy
day · upright
mug · heavy
dim · upright
glass · clear
day · upright
EDGE
glass · clear
dim · upright
glass · heavy
day · upright
EDGE
glass · heavy
dim · upright

Every option you add multiplies the suite. Choosing which combinations matter, marking the edge cases, and writing one scoring rule that holds across all of them is the design work.

What you get back

A written guide anyone can rerun next month.

Every suite ships with a scoring guide in plain language, examples for each call, and a version number, so a result from March and a result from June mean the same thing.

Shelf placement · scoring guide v1.1example guide
PASSItem upright on the target shelf, gripper released, within 60 seconds.
PARTIALItem on the shelf but tipped, or placed on the wrong level.
FAILItem dropped, a collision with the shelf, or not placed in 60 seconds.
EDGETransparent glass under dim or backlit light is scored, and reported separately.
NOTERecoveries within the attempt still count as PASS. Log them for the recovery rate.
What we design

The test cases and rules behind a number you can trust.

01
Scenario catalogue

The tasks your robot must handle, written as repeatable setups.

02
Variation plan

Objects, clutter, layout, lighting and start poses, chosen for coverage rather than count.

03
Edge cases

The combinations most likely to break the policy, scored and reported separately.

04
Scoring guide

Plain-language pass, partial and fail rules, with worked examples for each.

05
Setup & reset protocols

How each scene is arranged and reset, so trial 1 and trial 40 start the same way.

06
Versioning

Every change to the suite recorded with the reason, and old versions kept for history.

The pilot

From ad-hoc tests to a suite you can rerun.

1
Week 0

Pick one task family

Tell us the tasks that matter and send any footage or tests you already run.

2
Week 1

We draft the suite

Scenarios, variations, edge cases and a written scoring guide, agreed with your engineers.

3
Week 1-2

Score a first batch

We apply the guide to real rollouts, so unclear rules surface immediately.

4
Week 2

Hand over v1.0

The versioned protocol, the first scores, and our agreement numbers. You decide whether to scale.

FAQ

Questions, answered straight.

A fixed set of tasks, scene variations and scoring rules that a robot policy is tested against, so different model versions can be compared fairly. The value comes from the scenarios being representative and the scoring being written down and applied the same way every time.

Public benchmarks are useful for comparing against the field, but they rarely match your robot, your objects or your customers' environments. Most teams need an internal suite built around the tasks they actually ship.

We can score recorded trials against the suite remotely from day one. Running live trials, on your robots over a remote link or on a bench staged for your project, is scoped with you when you need it.

It is versioned. When your robot, objects or target tasks change, we propose updates, record what changed and why, and keep old versions so historical results stay interpretable.

With a small paid pilot on one task family: scenarios, variations and scoring guide drafted with your team, and a first batch scored against it. Pricing is scoped to the work rather than fixed.

Turn your ad-hoc robot tests into a suite you can rerun.

Start with a small paid pilot on one task family. We draft the scenarios, variations and scoring guide with your team, score a first batch against it, and hand over the protocol before you scale.