Services
Robotics data annotation, done by specialists

keypoints · grasps
Egocentric & Hand Labeling
First-person footage is how most humanoid and manipulation models learn to use their hands. We label all 21 keypoints per hand, grasp taxonomy, and the exact frames where a hand makes and breaks contact with an object.
- 21-point hand skeletons, per hand, per frame
- Grasp type (power, precision, pinch, lateral, custom taxonomy)
- Hand-object contact start and end frames
- Object boxes with consistent track IDs
- Deliverables
- Per-frame keypoint JSON · Contact event list (CSV) · QA report with reviewer notes
- Formats
- COCO keypoints · JSON · CSV · Custom schema
- Starting from
- Custom quote

steps · captions
Action Segmentation & Captions
We break long demonstrations into clean, timestamped steps and write dense captions a vision-language model can learn from. Step boundaries follow your verb list, so segments line up across the whole dataset.
- Start/end timestamps for every task step
- Verb-object labels from your taxonomy
- Dense natural-language captions per segment
- Episode-level task summaries
- Deliverables
- Segment file (JSON / CSV) · Caption file · Taxonomy notes for edge cases
- Formats
- JSON · CSV · WebVTT · LeRobot-style dataset · Custom schema
- Starting from
- Custom quote

success · failure
Robot Rollout Review
Policy evaluation needs more than a pass/fail flag. Reviewers watch each rollout, mark success or failure against your criteria, tag the failure mode, and timestamp the first frame where it went wrong.
- Success / partial / failure per episode
- Failure-mode tags (missed grasp, collision, drop, timeout…)
- Timestamp of first error
- Short free-text reviewer note
- Deliverables
- Per-episode review sheet · Failure-mode summary · Clips of flagged moments on request
- Formats
- CSV · JSON · Google Sheets
- Starting from
- Custom quote

boxes · masks · tracks
Image & Video Annotation
The fundamentals, done carefully: tight boxes, clean polygon edges, pixel masks, and object tracks that keep the same ID through occlusion.
- Bounding boxes and rotated boxes
- Polygons and semantic / instance masks
- Multi-object tracking with persistent IDs
- Attributes (occluded, truncated, state)
- Deliverables
- Label files in your format · Class list and guideline notes · QA report
- Formats
- COCO · YOLO · Pascal VOC · CVAT XML · JSON
- Starting from
- From $20 per 300 images

lidar · depth · multi-cam
3D & Multi-Sensor
When a robot sees the world through several sensors, labels have to agree in every one. We annotate 3D cuboids in point clouds and keep IDs and boxes consistent across depth and camera views.
- 3D cuboids with heading in point clouds
- Depth-map cleanup and masks
- Cross-camera ID association
- Sensor-frame consistency checks
- Deliverables
- 3D label files · Per-camera 2D projections · Consistency report
- Formats
- Depth maps (16-bit PNG) · 3D trajectories (JSON) · KITTI · Custom schema
- Starting from
- Custom quote

dedupe · anonymize · fix
Data Cleaning & QA
Sometimes the problem is the data you already have. We find duplicates and near-duplicates, blur faces and plates, convert between formats, and audit existing labels to fix the ones your model keeps tripping over.
- Duplicate and near-duplicate removal
- Face and licence-plate anonymization
- Format conversion between label schemas
- Label audits with corrected files
- Deliverables
- Cleaned dataset · Change log (what was removed or fixed, and why) · Audit summary
- Formats
- Any of COCO / YOLO / VOC / CVAT / JSON / CSV
- Starting from
- Custom quote
Large volumes
For 10,000+ items or a continuous pipeline, we price per item after the pilot, once we both know the real time per item. No minimum contract.
Security
Confidential handling by our own team, NDAs on request, face and plate anonymization offered, and data deleted after delivery.
See your own data labeled before you pay anything.
Send 10 to 20 frames or a short clip. We label them to your guidelines and send a quality report within 48 hours.