Skip to content
Annoroid
Request a free pilot
All work

Sample project

First-person photos: hands, forearms and what they're holding

EgocentricHand-object
First-person photos: hands, forearms and what they're holding — raw frame and labeled output
RAWLABELED
Data
Images · up to 1536×2048
Volume
3 images · 21 labeled instances
Status
Sample project

The problem

Egocentric data is what humanoid hands learn from, but hands at unusual angles, rings and watches, and objects half-hidden in the grip are easy to miss or mislabel.

What we did

  • 9 classes including hand, forearm, ring, smartwatch and held objects
  • Tight boxes on small worn items
  • Scene captions for each photo

Result

21 labeled instances across 3 first-person photos, exported as COCO, YOLO and CVAT.

All three photos, labeled

Photo 1 — raw frame and labeled output
RAWLABELED
Photo 2 — raw frame and labeled output
RAWLABELED
Photo 3 — raw frame and labeled output
RAWLABELED

Sample output

An excerpt from the delivered files. "source": "ai" marks boxes propagated by the segmentation model; every one was reviewed by a person before export.

ego.jsonexcerpt
{
  "image_id": 3,
  "category": "camera lens",
  "bbox": [384.96, 222.08, 194.88, 222.72],
  "attributes": {
    "track_id": "obj_9aba363ba103",
    "source": "manual"
  }
}

Photos (CC0 1.0): Alina Kakshapati; Faisal Ahammad; Jonas Svidras. Labels and captions by Annoroid.

Request a free pilot for your own data.

Send 10 to 20 frames or a short clip. We label them to your guidelines and send a quality report within 48 hours.