Your success rate dropped. Here's exactly why.
A dedicated team watches every failed and borderline rollout, tags the cause against a taxonomy agreed with your engineers, marks where it began and whether the robot recovered, and hands you a weekly breakdown of where the policy breaks.
- Independent & neutral
- Your taxonomy, written down
- Remote, on recorded rollouts
- Never crowdsourced
One failed rollout: events marked on the timeline, then tagged by cause, start point and recovery.
A number shows a drop, not a cause
Pass rate fell four points. Grip, perception, planning or recovery? Without tags, engineers rewatch hours of video to find out.
The obvious tag is often wrong
A plate that hit the rack and then fell was a collision, not a drop. Tag it wrong and the team fixes the gripper, which was fine.
Tags drift without a written taxonomy
If 'missed grasp' means something different each week, version-to-version comparisons stop meaning anything.
Tag three failures. Then check the taxonomy.
Each is a failed rollout from a household robot. Pick the cause, then see what the taxonomy says and why the obvious answer can send engineers to fix the wrong thing.
Load a plate into the dish rack
- 0:02grasp plate edge
- 0:05rotate to vertical
- 0:07plate rim hits rack tine
- 0:08plate knocked out of grip
Hand the red cup to the person
- 0:01reach
- 0:03grasp blue cup
- 0:06lift cleanly
- 0:09hand over
Open the top drawer
- 0:02reach handle
- 0:03grip handle
- 0:04pull, drawer stuck
- 0:05pull again, same force
- 0:60timeout
Tag all three to compare with the taxonomy.
Failure modes ranked, with the fix they point to.
Every tagged rollout rolls up into one view per model version: which causes dominate, how that moved since last week, and which part of the stack each one points at.
The failure detail a success rate hides.
From a taxonomy written with your engineers, with examples for every category.
The step and moment the failure actually started, not where the robot gave up.
Whether the robot recovered on its own, how, and how long it took.
Every time a human operator had to step in, and why.
Borderline successes that pass today and are likely to fail tomorrow.
Failures that fit no category, flagged so the taxonomy grows deliberately.
From a pile of failures to a ranked list.
You send failures
Recent failed and borderline rollouts from one policy, with whatever notes your team already keeps.
We draft the taxonomy
Categories built from your own footage, with examples, agreed with your engineers before tagging starts.
Tag and double-check
Every rollout tagged. A sample of every batch is tagged twice to measure agreement.
Breakdown and agreement
Failure modes ranked, what each points to, and our agreement numbers. You decide whether to scale.
Questions, answered straight.
Robot failure analysis reviews the attempts a robot policy got wrong and records why: a missed grasp, a dropped object, the wrong object, a collision, getting stuck. Rolled up across many rollouts, it shows engineers which failure modes to fix first.
You do, with us. We draft a failure taxonomy from a sample of your rollouts, agree it with your engineers, and write it down with examples. The team then tags against that document, and we propose changes when new failure types appear.
Vision-language models can draft a description, and we use them where they help. Deciding the actual cause, the moment it began and whether the robot recovered is still a judgment call, and that is what our evaluators own.
Every rollout tagged with cause, step and recovery, plus a weekly breakdown of failure modes by task and by model version, and the agreement rate between our evaluators on a double-scored sample.
With a small paid pilot on recent failed rollouts from one policy. Pricing depends on volume and detail, so we scope it with you. The pilot shows you the output before you commit to more.
Send us your failed rollouts. Get back the reasons.
Start with a small paid pilot on one policy's recent failures. We draft the failure taxonomy with your team, tag every rollout, and show you the breakdown and our agreement numbers before you scale.