Skip to content
annoroid
All samples

Live sample

03

Preference Data for Reward Models

“We need preference data for reward models and don't have trained raters for robot behavior.”
What we do
Pairwise and ranked judgments of robot behavior on smoothness, safety and efficiency, from raters trained on manipulation.
What you receive
Preference pairs with rater agreement per criterion, ready for reward modeling.

Interactive sample · watch two real rollouts, pick one, set the criteria

A
B
00:00.0

Footage: ALOHA static coffee, LeRobot / ALOHA project (MIT), two episodes of the same task. Ranking: yours.

Which rollout is better?

All samples are built on public data: robot episodes from the open ALOHA dataset (LeRobot, MIT licence) and public-domain kitchen footage from the USDA, processed by us. Samples marked “Illustrative” use example data. Sources

Other samples

Want this on your own robot data? Start with a pilot.

Judge the output yourself. The pilot is free, results come back in 48 hours, and there's no contract.