WeTalkRobots · About

LARGE VIDEO PLANNER ENABLES GENERALIZABLE ROBOT CONTROL

Learning from Humans · 17/12/2025

Technical Analysis Large Video Planner "LVP" is a video foundation model for robotics that generates video plans, which we retarget into executable robot motions. High-level Analogy: Imagine you want a robot to perform a new chore, like 'open the first aid box' or 'tear off the tape.' Instead of trying to teach the robot with complex instructions or showing it exactly how to move its joints, you show it a video of a human doing that exact task. The robot watches the video, figures out the general 'flow' of the human's movement – where their hand goes, how they grasp, how they manipulate the object – and then translates that visual plan into its own robot movements. It's like a chef watching a cooking show to learn a new recipe, then adapting the human chef's techniques to their own kitchen tools and ingredients. Motivation of the Work Developing robots that can handle diverse tasks in new environments is a major challenge. Recently, robot foundation models have emerged, often by enhancing powerful language and image models (called MLLMs) to also output actions. These are known as Vision-Language-Action (VLA) models. The idea is that the vast knowledge MLLMs gained from analyzing internet text and images could teach robots how to act. However, there's a big hurdle: robot action data is extremely scarce compared to the massive amounts of text and image data available online. This means VLA models often rely on an 'asymmetric' transfer, where pre-trained MLLM knowledge is only…

Continue to interactive post