WeTalkRobots · About

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

World Action Models · 19/12/2025

Technical Analysis "Mimic-Video" is a Video-Action Model (VAM) that pairs a pretrained Internet-scale video model with a flow matching based action decoder conditioned on its latent representations. The decoder serves as an Inverse Dynamics Model (IDM), generating low-level robot actions from the latent representation of video space action plans. High-level Analogy: Imagine you want to teach a robot to cook. Traditional robots (VLAs) are like students who have only read recipe books (images and text). They know what a carrot is and what 'chop' means, but they've never seen anyone actually do it. So, when they try to chop carrots, they have to figure out all the physics of gripping, cutting, and object interaction from scratch, based only on a few hands-on demonstrations. This is very slow and requires many expensive demonstrations. mimic-video robots are like students who have watched thousands of cooking videos on YouTube. They've seen people chop carrots, stir pots, and flip pancakes countless times. They have an intuitive understanding of how things move, deform, and interact (visual dynamics and physical causality). When you then show them just a few demonstrations of chopping carrots with a robot arm, they don't need to learn the basic physics. Instead, their 'video brain' quickly generates a 'mental video plan' of how the chopping should look. Their 'action brain' then simply translates this visual plan into the precise robot arm movements. Because they already…

Continue to interactive post