01
Policy Evaluation and Failure Analysis
“We generate more robot rollouts than we can review.”
- What we do
- Robotics reviewers judge real and sim rollouts: success or fail, the moment it failed, why, and a quality score. The result feeds your evals, reward models and RL post-training.
- What you receive
- Scored rollouts in your schema, plus a failure-cause report that tells you what to fix next.
Interactive sample · switch raw / evaluated on a real rollout, or see a failure diagnosed
Subtask 1/7 · Pick up the capsule0 of 7 checkpoints passed
00:00.0
OutcomeIn progress
Quality score
Footage: ALOHA static coffee, LeRobot / ALOHA project (MIT). Evaluation: Annoroid.