WeTalkRobots · About

Human Egocentric Data: How to Extract Annotations

Reports · 17/09/2026

Seeing the World Through Human Eyes: Collecting and Labeling Egocentric Data for Robot Learning From wrist-mounted cameras to physics-aware force annotations — a look at the methods behind the egocentric data boom. --- Robots and AI systems that learn from humans face a fundamental challenge: most of what humans do happens from a first-person perspective. We cook, assemble, grasp, and gesture — all while our hands occupy the center of our visual field. Egocentric video, captured from head- or body-mounted cameras, is rapidly becoming one of the most valuable data modalities for robot imitation learning, AR/VR, and embodied AI. But raw video is not enough. To train useful models, researchers need structured supervision — camera trajectories, 3D hand poses, contact regions, interaction forces, and language labels — all precisely aligned to each frame. This post surveys the key trends in how that data is being collected and annotated, drawing on five recent works: WiLoR, Dyn-HaMR, HaWoR, MINT, and EgoPHI. --- The Egocentric Data Landscape Over the past few years, several large-scale egocentric datasets have become the community's shared testbeds: Ego4D — thousands of hours of daily activity from hundreds of participants worldwide EPIC-KITCHENS — fine-grained kitchen interaction video with action labels EgoDex — dexterous manipulation footage for robot imitation learning HOT3D — egocentric video with synchronized ground-truth 3D hand and object poses from head-mounted devices…

Continue to interactive post