Test the robot you'll ship, not the demo you filmed.
We design the scenarios, variations and edge cases your policy should face, write a scoring guide plain enough to apply the same way every time, and keep it versioned as your robot and your product change.
- Independent & neutral
- Written, versioned protocols
- Built with your engineers
- Never crowdsourced
One base task fanned out across object, clutter and lighting, with edge cases marked and the scoring guide written.
Demo setups flatter policies
A policy tuned on three favourite scenes can look strong there and fall apart on the first unfamiliar kitchen.
Variation is a design choice
Every object, clutter level and light condition multiplies the suite. Picking the combinations that matter is the work.
Rules in someone's head don't rerun
If the pass criteria aren't written down with examples, results change when the person running the test changes.
Build a test suite. Watch it grow.
Pick the variations for one household task. See how fast the suite multiplies, which combinations become edge cases, and how many trials it takes to score it.
day · upright
dim · upright
day · upright
dim · upright
day · upright
dim · upright
day · upright
dim · upright
Every option you add multiplies the suite. Choosing which combinations matter, marking the edge cases, and writing one scoring rule that holds across all of them is the design work.
A written guide anyone can rerun next month.
Every suite ships with a scoring guide in plain language, examples for each call, and a version number, so a result from March and a result from June mean the same thing.
The test cases and rules behind a number you can trust.
The tasks your robot must handle, written as repeatable setups.
Objects, clutter, layout, lighting and start poses, chosen for coverage rather than count.
The combinations most likely to break the policy, scored and reported separately.
Plain-language pass, partial and fail rules, with worked examples for each.
How each scene is arranged and reset, so trial 1 and trial 40 start the same way.
Every change to the suite recorded with the reason, and old versions kept for history.
From ad-hoc tests to a suite you can rerun.
Pick one task family
Tell us the tasks that matter and send any footage or tests you already run.
We draft the suite
Scenarios, variations, edge cases and a written scoring guide, agreed with your engineers.
Score a first batch
We apply the guide to real rollouts, so unclear rules surface immediately.
Hand over v1.0
The versioned protocol, the first scores, and our agreement numbers. You decide whether to scale.
Questions, answered straight.
A fixed set of tasks, scene variations and scoring rules that a robot policy is tested against, so different model versions can be compared fairly. The value comes from the scenarios being representative and the scoring being written down and applied the same way every time.
Public benchmarks are useful for comparing against the field, but they rarely match your robot, your objects or your customers' environments. Most teams need an internal suite built around the tasks they actually ship.
We can score recorded trials against the suite remotely from day one. Running live trials, on your robots over a remote link or on a bench staged for your project, is scoped with you when you need it.
It is versioned. When your robot, objects or target tasks change, we propose updates, record what changed and why, and keep old versions so historical results stay interpretable.
With a small paid pilot on one task family: scenarios, variations and scoring guide drafted with your team, and a first batch scored against it. Pricing is scoped to the work rather than fixed.
Turn your ad-hoc robot tests into a suite you can rerun.
Start with a small paid pilot on one task family. We draft the scenarios, variations and scoring guide with your team, score a first batch against it, and hand over the protocol before you scale.