200,000+ Hours of Egocentric First-Person Video Corpus for Humanoid Systems
Case Studies
Video

200,000+ Hours of Egocentric First-Person Video Corpus for Humanoid Systems

Humanoid robots don’t learn dexterity from staged demonstrations. They learn from observing how humans move through real environments, how hands reach, adjust, and recover without explicit instruction. That signal exists only in the environments robots are expected to operate in: homes. Not labs, not studios, not simulations.

We built an egocentric video corpus of up to 200,000 hours, captured from head-mounted smartphones while participants performed everyday household tasks in their own homes. The dataset provides continuous, first-person views of intentional human motion, the exact perspective humanoid systems require for task learning.

Project Overview

CategorySpecification
ModelHumanoid robotics, egocentric motion, and task learning
VolumeUp to 200,000 hours of video
Capture MethodHead-mounted smartphones, egocentric, first-person perspective
EnvironmentReal residential homes, apartments, and houses
TasksNatural household chores are performed during normal routines
UploadAutomated via client collection application over WiFi / 5G
Key RequirementHands visible; continuous, purposeful task execution

Egocentric Motion at Scale needs Physical Infrastructure

The dataset required participants to record from eye level while performing real chores in their own homes. This introduced constraints that standard collection methods couldn’t meet.

The recording setup, a head-mounted smartphone, had to become routine enough that participants behaved naturally. Tasks couldn’t be staged or performed for the camera. Environments couldn’t be controlled. Kitchens differed in layout, lighting, and spatial constraints. That variability wasn’t noise. It was the signal.

Quality requirements were equally strict. Hands needed to remain visible. Tasks needed to be intentional and continuous. Privacy violations, faces, children, or identifiable information, had to be excluded at capture, not after collection. At this scale, post-hoc filtering would have made the dataset slow and inefficient to build.

Maintaining technical fidelity across thousands of distributed recording hours required continuous feedback during collection, clear recording protocols, and infrastructure that could catch problems before they propagated.

200,000 Hours of Egocentric Motion, Captured Through Distributed Infrastructure in 3 Months

Scaling to 200,000 hours so quickly required removing hardware as a bottleneck. Participants procured compatible headstraps and smartphones locally and were reimbursed after verification, allowing collection to scale across regions without centralized device deployment.

Data was captured by domestic blue-collar workers across South Asia, Southeast Asia, and the Middle East, recording chores they already performed, cooking, cleaning, washing clothes, folding garments, and maintaining living spaces. Because recording was embedded into real routines, motion remained natural and unscripted. A distributed validation layer ensured capture quality and protocol adherence, preventing unusable or non-compliant footage from entering the pipeline.

The result is a large-scale egocentric video corpus built through participant-led infrastructure, reflecting how tasks are actually performed in real homes, and providing training data that humanoid systems can generalize from.