Sample project
Tomato salad prep: tracking through occlusion, with hand-object relationships


- Data
- Video · 1920×1080
- Volume
- 150 frames · 2,526 labeled instances
- Status
- Sample project
The problem
Small objects like garlic cloves appear, split and disappear into a bowl. Tracks that lose their ID or keep a box after the object is gone teach a model the wrong thing.
What we did
- 19 object classes tracked with persistent IDs (21 tracks)
- Tracks ended exactly where an object leaves the shot or is combined into another
- 20 hand-object contacts with the object each hand is touching
- 9 action captions timed to the frame, plus scene captions
- Every track reviewed on a contact sheet before export
Result
150 frames with 2,526 labeled instances, 21 tracks and 20 hand-object contacts.
Labeled video · 30 seconds
- Caption track
- 00:00The last tomato half is sliced and the slices are added to the glass bowl.
- 00:01The board is cleared and a garlic bulb is brought over.
- 00:04Garlic cloves are broken off the bulb and peeled by hand.
- 00:12The peeled garlic is minced with the knife.
- 00:18The minced garlic is gathered by hand and sprinkled over the tomatoes.
- 00:22Salt is pinched from the salt bowl and sprinkled over the tomatoes.
- 00:26Pepper is pinched from the pepper bowl and sprinkled over the tomatoes.
- 00:28Serving tongs are brought in and the bowl is drawn to the centre of the board.
- 00:29The salad is tossed with the tongs.
Frames, depth and review sheet





Sample output
An excerpt from the delivered files. "source": "ai" marks boxes propagated by the segmentation model; every one was reviewed by a person before export.
{
"start_ms": 3604,
"end_ms": 12212,
"kind": "action",
"text": "Garlic cloves are broken off the bulb and peeled by hand.",
"source": "manual"
}Source footage: “Farmers Market Food Demo - Spanish Tomato Salad”, US Department of Agriculture (Agricultural Marketing Service), via Wikimedia Commons (Public domain). Labels, captions, tracks and depth by Annoroid.