03
Preference Data for Reward Models
“We need preference data for reward models and don't have trained raters for robot behavior.”
- What we do
- Pairwise and ranked judgments of robot behavior on smoothness, safety and efficiency, from raters trained on manipulation.
- What you receive
- Preference pairs with rater agreement per criterion, ready for reward modeling.
Interactive sample · watch two real rollouts, pick one, set the criteria
A
B
00:00.0
Footage: ALOHA static coffee, LeRobot / ALOHA project (MIT), two episodes of the same task. Ranking: yours.
Which rollout is better?