WeTalkRobots · About

Masquerade: Learning from In-the-wild Human Videos using Data-Editing

Learning from Humans · 12/08/2025

Technical Analysis "Masquerade" is a pipeline that turns human videos into “robotized” demonstrations by (i) estimating 3-D hand poses, (ii) inpainting the human arms, and (iii) overlaying a rendered bimanual robot that tracks the recovered end-effector trajectories. High-level Analogy: Imagine you want to teach a robot how to perform various tasks by showing it videos of humans doing them (like cooking from a YouTube tutorial). The big problem is that humans and robots look very different – a human hand is not a robot gripper. It's like trying to teach a young child how to draw by showing them complex blueprints. The child might get some ideas, but the visual styles are so different that it's hard to directly copy. Masquerade tackles this by acting as a 'visual translator' or 'robot costume designer'. First, it takes countless human videos and: Analyzes human movements: It figures out exactly where the human's hands are and how they move. Digitally removes humans: It then erases the human arms from the video. Dresses a robot in 'human' movements: Crucially, it overlays a rendered robot's arms and grippers onto the video, making them follow the exact same path the human's hands originally took. Now, the robot in the video is 'masquerading' as the human. Then, this 'robotized' video data, showing robots performing tasks, is used to teach a real robot. The robot also gets a few real-life demonstrations, but it keeps 'watching' the vast library of 'robotized' human videos at…

Continue to interactive post