Skip to content
annoroid

Pipeline

The pipeline

Software does the first pass. People verify every episode. You get the output and the reasons behind it.
  1. 01

    Ingest

    We read your data as it is.

    • Read from your storage (AWS S3, Google Cloud Storage, Cloudflare R2, any S3-compatible storage, Hugging Face Hub) or from direct upload
    • Parse the native format: LeRobot, HDF5, ROS 2 bags / MCAP, MP4 + IMU
    • Build an episode index: streams, rates, durations, camera count
  2. 02

    Automated checks

    Software runs on every frame.

    • Timestamp gaps and dropped frames, against your expected rate
    • Stream alignment: cameras vs. joint state vs. actions
    • Frozen sensors (identical consecutive readings)
    • Idle time at episode start and end
    • Episode boundaries and done flags
    • Commanded vs. measured tracking (grippers, joints)
    • Video: blur, occlusion, exposure, missing frames
    • Egocentric: IMU and video sync, consent and provenance metadata present

    Output · A per-episode check report, plus proposed segments, labels and rejections.

  3. 03

    Human verification

    People check every episode, not a sample.

    • Trained reviewers confirm or correct segment boundaries, task descriptions, success or failure and failure type
    • A second reviewer re-checks a blind sample of each batch to measure accuracy
    • Disagreements go to a lead reviewer, and the guideline is updated so the next batch is consistent
  4. 04

    Outputs

    Your data back, plus everything we learned about it.

    • Cleaned episodes in your original format (originals untouched)
    • quality: score and individual checks per episode (JSON / Parquet)
    • rejections: episode IDs with reason codes
    • segments: start and end frames and step names per episode
    • language: task instruction plus per-segment instructions
    • failures: outcome, failure frame and failure type
    • A delivery report: what we did, what we found, measured accuracy

Before / after

Example: a public LeRobot dataset

Source: lerobot/aloha_static_coffee (MIT). 50 episodes · 55,000 frames · 4 cameras · 14-joint state, action and effort at 50 fps. Every number below comes from running our automated checks on the dataset's own files.

Before

  • One task sentence for all 50 episodes
  • No step boundaries, contact events or per-step instructions
  • No quality signal per episode

“Place the coffee capsule inside the capsule container, then place the cup onto the center of the cup tray, then push the 'Hot Water' and 'Travel Mug' buttons.”

Automated checks found

  • Timestamp gaps (expected 20 ms steps)0 gaps in 55,000 frames
  • Episode boundaries and done flags50 of 50 correct
  • Camera video vs joint data alignment200 of 200 camera segments aligned
  • Frozen sensor frames (identical consecutive state)63 repeated readings across 30 episodes; at most 5 identical in a row (0.08 s)
  • Idle time at episode ends5.1 s in total; longest ep 31 (0.78 s)
  • Gripper tracking (commanded vs measured)Left gripper differs 0.038–0.116 rad vs 0.007–0.038 on the right; left is over 2× the right in 43 of 50 episodes
  • Label depthOne task sentence for all 50 episodes; no step boundaries, contacts or per-step instructions

After: what a delivery adds

  • Seven step boundaries per episode, from pick up the capsule to pressing the buttons
  • Gripper contact and release events for the capsule and the cup
  • Low-level instructions per step under the dataset's own task sentence
  • A calibration note on the left gripper, with the episodes most affected

Clean capture. The gap is label depth, and one gripper worth a calibration check.

Real output: episode 30 check record

{
  "episode": 30,
  "frames": 1100,
  "timestamp_gaps": 0,
  "max_step_s": 0.02,
  "frame_index_gaps": 0,
  "nan_state": 0,
  "frozen_frames": 2,
  "longest_frozen_run": 1,
  "state_spikes": 0,
  "nan_action": 0,
  "idle_head_s": 0,
  "idle_tail_s": 0.14,
  "tracking_err": {
    "left_waist": 0.0046,
    "left_shoulder": 0.0194,
    "left_elbow": 0.0202,
    "left_forearm_roll": 0.0054,
    "left_wrist_angle": 0.0105,
    "left_wrist_rotate": 0.004,
    "left_gripper": 0.1157,
    "right_waist": 0.0144,
    "right_shoulder": 0.024,
    "right_elbow": 0.0245,
    "right_forearm_roll": 0.0051,
    "right_wrist_angle": 0.0132,
    "right_wrist_rotate": 0.0048,
    "right_gripper": 0.0351
  },
  "done_flags": 1,
  "last_frame_done": true,
  "cams_checked": 4,
  "cams_misaligned": [],
  "tasks": [
    "Place the coffee capsule inside the capsule container, then place the cup onto the center of the cup tray, then push the 'Hot Water' and 'Travel Mug' buttons."
  ]
}

Download the sample output (zip, 8 KB): this record and four more as JSON and Parquet, plus the quality report for all 50 episodes.

Run this on your own data.

Send 100 episodes. You get the cleaned dataset, the check report and the labels back in 48 hours, free.