Collect
Trained recorders capture task episodes on calibrated head-mounted rigs. Every session starts with a signed consent form and a scene check.
We record how people actually do things — hands, gaze, tools, mess and all — then parse and filter that footage into training-ready episodes for embodied AI.
Why first-person
A robot arm never sees the scene from across the room. It sees what a hand sees: an object half-occluded by fingers, a handle at an odd angle, a surface that shifts as the head turns.
That viewpoint gap is why models trained on scraped web video stall the moment they touch real hardware. KarYeah closes it — recording from the body, keeping only what a policy can learn from.
What "parsed" means
380+ recorders · North India
The pipeline · in order
Trained recorders capture task episodes on calibrated head-mounted rigs. Every session starts with a signed consent form and a scene check.
Footage is synced, de-warped and cut into action segments with hand pose, gaze, object tracks and a plain-language narration per segment.
Blurred frames, dropped tracking, bystander faces and near-duplicate takes are removed. What survives is rare, clean and genuinely varied.
Episodes are packed into LeRobot, RLDS or your schema, with a data sheet listing provenance, consent status and known gaps.
What survives the filter
Figures updated quarterly · last revision Q2 2026
Ready to license
Prep, pouring, plating and clean-up across home kitchens, dhabas and canteen lines. Dense two-hand manipulation with constant occlusion.
Two-wheeler service, electrical repair and carpentry. Tool selection, fine alignment, and the failure-and-retry loops policies rarely see.
Picking, stacking, barcode handling and shelf resets in small shops and micro-fulfilment floors. Heavy on navigation plus grasp.
Consent & provenance
Recorders are paid, briefed and free to pull a session back after the fact. Buyers get the paperwork, not just the pixels.
Read the data charterQuestions we get
Egocentric data is recorded from a person's own point of view — usually head-mounted video, plus gaze, hand pose, audio and motion sensors. Because the camera moves with the body, it captures the same viewpoint a robot or wearable agent has to act from.
Every recorder signs a consent agreement before capture. Faces, licence plates, screens and documents are blurred during the parse stage, and a human reviewer checks a sample of every batch before release. Recorders can withdraw a session for 30 days after it is filed.
Video as MP4 or MKV with per-frame timestamps, annotations as JSON or Parquet, and full episodes in LeRobot, RLDS or a schema you specify. Delivery by S3, GCS or encrypted physical drive.
Yes. Commissioned collection is most of our work. Send the task list and the rig constraints; a scoped pilot of 40–100 episodes usually runs in three to four weeks.
KarYeah is a venture of 172 Tech, built and run from Chandigarh. The name is Punjabi shorthand — kar, to do. The site was designed by 530 Expert.