WeTalkRobots · About

Phantom: Training Robots Without Robots Using Only Human Videos

Learning from Humans · 02/03/2025

Technical Analysis "Phantom" is a scalable framework for training manipulation policies directly from human video demonstrations, requiring no robot data. It converts human demonstrations into robot-compatible observation-action pairs using hand pose estimation and visual data editing, by inpainting the human arm and overlaying a rendered robot to align the visual domains. High-level Analogy: Imagine you want to teach a robot to perform various tasks like picking up objects, tidying up, or tying knots. Traditionally, this is like trying to teach a child a new skill by physically moving their hands – it's slow, requires constant supervision, and you can only do it with one child at a time in one place. You also need to physically demonstrate the exact movements. This research paper, 'Phantom,' proposes a much more scalable way, like teaching a robot by showing it countless videos of people performing these tasks. But there are two main hurdles: Humans look and move differently than robots: A human hand is soft and flexible, a robot gripper is hard and mechanical. The robot needs to 'see' itself doing the task, not a human. Human videos don't include robot commands: A video shows what a human does, but not the precise motor instructions (like 'move gripper to X, Y, Z coordinates and open by this much') that a robot needs. Phantom acts like a super-smart 'virtual body double' and 'motion translator' for the robot: Motion Translator: It watches the human videos very carefully,…

Continue to interactive post