Skip to content
annoroid

Samples

Try the work on real robot data

Six live samples, one per core service. Evaluate a rollout, rank two policies, repair a dataset: in the browser, on real footage, with the output we deliver.

All samples are built on public data: robot episodes from the open ALOHA dataset (LeRobot, MIT licence) and public-domain kitchen footage from the USDA, processed by us. Samples marked “Illustrative” use example data. Sources

01

Policy Evaluation and Failure Analysis

“We generate more robot rollouts than we can review.”
What we do
Robotics reviewers judge real and sim rollouts: success or fail, the moment it failed, why, and a quality score. The result feeds your evals, reward models and RL post-training.
What you receive
Scored rollouts in your schema, plus a failure-cause report that tells you what to fix next.

Interactive sample · switch raw / evaluated on a real rollout, or see a failure diagnosed

Subtask 1/7 · Pick up the capsule0 of 7 checkpoints passed
Evaluated
00:00.0
OutcomeIn progress
Quality score

Footage: ALOHA static coffee, LeRobot / ALOHA project (MIT). Evaluation: Annoroid.

02

Episode Structuring for Robot Learning

“Our teleop data is hours of video with no structure a policy can learn from.”
What we do
Teleop and egocentric video turned into training-ready episodes: subtasks, contact events, grasp phases, hand and object state, and language instructions at several levels.
What you receive
Structured episodes in JSON, CSV or LeRobot-compatible formats, aligned to your timestamps.

Interactive sample · compare raw and structured footage live, click a Hold

20 objects · 2,711 instances · 23 contacts · 12 captions
Raw
Structured by Annoroid
00:00.0

Now · action

Yellow watermelon wedges are cut into cubes on the cutting board.

Now · hands in contact

L · knife

Instruction level
Subtask
Contacttop L · bottom R
Grasptop L · bottom R
Instruction

Hold · high-level

“Make a watermelon, cucumber and basil salad”

Low-level

Click a Hold in the Grasp track

Click a green “Hold” in the Grasp track to expand it into instructions. Footage: USDA (public domain). Labels: Annoroid.

03

Preference Data for Reward Models

“We need preference data for reward models and don't have trained raters for robot behavior.”
What we do
Pairwise and ranked judgments of robot behavior on smoothness, safety and efficiency, from raters trained on manipulation.
What you receive
Preference pairs with rater agreement per criterion, ready for reward modeling.

Interactive sample · watch two real rollouts, pick one, set the criteria

A
B
00:00.0

Footage: ALOHA static coffee, LeRobot / ALOHA project (MIT), two episodes of the same task. Ranking: yours.

Which rollout is better?

04

Dataset Audit and Repair

“Our research engineers are spending their time fixing data instead of training models.”
What we do
We measure what's wrong in an existing dataset (empty, wrong, inconsistent or duplicate data), repair it, and tell you which sources cause it.
What you receive
The repaired dataset plus an audit report with error rates by source and type.

Interactive sample · before and after on a real repair run of 900 frames

Clean-up run · 3 raw kitchen clips · 900 frames

478 kept332 near-duplicate90 blur / transition
frame 01-f0001brightness 181.3sharpness 209.3verdictblur / transition

Real run on our raw clips: 478 of 900 frames worth training on. Hover a tile to read its log row. See our audit of a public LeRobot dataset →

05

Cross-Embodiment Data Standards

“Data from different robots and sources doesn't fit one schema.”
What we do
One taxonomy and format across datasets, robots and sites, so every new source plugs into training without a rewrite.
What you receive
A documented ontology, written guidelines and mapping scripts you own.

Interactive sample · hover a raw label to see where it goes

Before · labels as they arrivedIllustrative

15 labels, three names for the same cup, no structure.

After · one hierarchy

  • object
    • container
      • cup3
    • appliance
      • kettle1
  • action
    • grasp
      • power2
      • precision1
    • pour1
  • agent
    • hand
      • left2
      • right1
  • outcome
    • success1
    • failure
      • drop2
      • slip1

10 leaf classes with definitions, edge-case notes and a version number.

06

Managed Data Pipeline with SLA

“Our data work can't keep up with how fast we collect data.”
What we do
A dedicated robotics team that runs evaluation and data preparation as an ongoing pipeline, at a committed turnaround.
What you receive
A weekly delivery cadence, with throughput and quality reporting.

Interactive sample · watch batches move through the pipeline

Delivery pipeline · dedicated robotics team

IllustrativeTurnaround agreed per order, in writing
Weekly delivery W1 W2 W3 W4 W5 W6Throughput and quality report with every delivery

Want this on your own robot data? Start with a pilot.

Judge the output yourself. The pilot is free, results come back in 48 hours, and there's no contract.