Sample project
First-person photos: hands, forearms and what they're holding
EgocentricHand-object


- Data
- Images · up to 1536×2048
- Volume
- 3 images · 21 labeled instances
- Status
- Sample project
The problem
Egocentric data is what humanoid hands learn from, but hands at unusual angles, rings and watches, and objects half-hidden in the grip are easy to miss or mislabel.
What we did
- 9 classes including hand, forearm, ring, smartwatch and held objects
- Tight boxes on small worn items
- Scene captions for each photo
Result
21 labeled instances across 3 first-person photos, exported as COCO, YOLO and CVAT.
All three photos, labeled






Sample output
An excerpt from the delivered files. "source": "ai" marks boxes propagated by the segmentation model; every one was reviewed by a person before export.
ego.jsonexcerpt
{
"image_id": 3,
"category": "camera lens",
"bbox": [384.96, 222.08, 194.88, 222.72],
"attributes": {
"track_id": "obj_9aba363ba103",
"source": "manual"
}
}Photos (CC0 1.0): Alina Kakshapati; Faisal Ahammad; Jonas Svidras. Labels and captions by Annoroid.