Robot failure analysis

Your success rate dropped. Here's exactly why.

A dedicated team watches every failed and borderline rollout, tags the cause against a taxonomy agreed with your engineers, marks where it began and whether the robot recovered, and hands you a weekly breakdown of where the policy breaks.

Tag a failure yourself ↓
  • Independent & neutral
  • Your taxonomy, written down
  • Remote, on recorded rollouts
  • Never crowdsourced

One failed rollout: events marked on the timeline, then tagged by cause, start point and recovery.

01

A number shows a drop, not a cause

Pass rate fell four points. Grip, perception, planning or recovery? Without tags, engineers rewatch hours of video to find out.

02

The obvious tag is often wrong

A plate that hit the rack and then fell was a collision, not a drop. Tag it wrong and the team fixes the gripper, which was fine.

03

Tags drift without a written taxonomy

If 'missed grasp' means something different each week, version-to-version comparisons stop meaning anything.

Try it

Tag three failures. Then check the taxonomy.

Each is a failed rollout from a household robot. Pick the cause, then see what the taxonomy says and why the obvious answer can send engineers to fix the wrong thing.

Failed rollout 1fail

Load a plate into the dish rack

  1. 0:02grasp plate edge
  2. 0:05rotate to vertical
  3. 0:07plate rim hits rack tine
  4. 0:08plate knocked out of grip
Your tag: what caused it?
Failed rollout 2fail

Hand the red cup to the person

  1. 0:01reach
  2. 0:03grasp blue cup
  3. 0:06lift cleanly
  4. 0:09hand over
Your tag: what caused it?
Failed rollout 3fail

Open the top drawer

  1. 0:02reach handle
  2. 0:03grip handle
  3. 0:04pull, drawer stuck
  4. 0:05pull again, same force
  5. 0:60timeout
Your tag: what caused it?

Tag all three to compare with the taxonomy.

What you get back

Failure modes ranked, with the fix they point to.

Every tagged rollout rolls up into one view per model version: which causes dominate, how that moved since last week, and which part of the stack each one points at.

Policy v12 · failure breakdownexample layout, not client data
Missed grasp
points to: grasp pose / perception
Slip or dropped
points to: grip force / compliance
Wrong object or target
points to: grounding / instruction following
Collision
points to: motion planning / clearance
Stuck, no strategy change
points to: recovery behaviour
What we tag

The failure detail a success rate hides.

01
Cause

From a taxonomy written with your engineers, with examples for every category.

02
Where it began

The step and moment the failure actually started, not where the robot gave up.

03
Recovery

Whether the robot recovered on its own, how, and how long it took.

04
Interventions

Every time a human operator had to step in, and why.

05
Near-misses

Borderline successes that pass today and are likely to fail tomorrow.

06
New failure types

Failures that fit no category, flagged so the taxonomy grows deliberately.

The pilot

From a pile of failures to a ranked list.

1
Week 0

You send failures

Recent failed and borderline rollouts from one policy, with whatever notes your team already keeps.

2
Week 1

We draft the taxonomy

Categories built from your own footage, with examples, agreed with your engineers before tagging starts.

3
Week 1-2

Tag and double-check

Every rollout tagged. A sample of every batch is tagged twice to measure agreement.

4
Week 2

Breakdown and agreement

Failure modes ranked, what each points to, and our agreement numbers. You decide whether to scale.

FAQ

Questions, answered straight.

Robot failure analysis reviews the attempts a robot policy got wrong and records why: a missed grasp, a dropped object, the wrong object, a collision, getting stuck. Rolled up across many rollouts, it shows engineers which failure modes to fix first.

You do, with us. We draft a failure taxonomy from a sample of your rollouts, agree it with your engineers, and write it down with examples. The team then tags against that document, and we propose changes when new failure types appear.

Vision-language models can draft a description, and we use them where they help. Deciding the actual cause, the moment it began and whether the robot recovered is still a judgment call, and that is what our evaluators own.

Every rollout tagged with cause, step and recovery, plus a weekly breakdown of failure modes by task and by model version, and the agreement rate between our evaluators on a double-scored sample.

With a small paid pilot on recent failed rollouts from one policy. Pricing depends on volume and detail, so we scope it with you. The pilot shows you the output before you commit to more.

Send us your failed rollouts. Get back the reasons.

Start with a small paid pilot on one policy's recent failures. We draft the failure taxonomy with your team, tag every rollout, and show you the breakdown and our agreement numbers before you scale.