WeTalkRobots · About

World Action Models are Zero-shot Policies

World Action Models · 17/02/2026

Technical Analysis "DreamZero"is a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by jointly predicting future world states and actions, using video as a dense representation of how the world evolves. By jointly modeling video and action, DreamZero learns diverse skills effectively from heterogeneous robot data without relying on repetitive demonstrations. High-level Analogy: Imagine you want to teach a robot to perform a wide variety of tasks, like setting a table, folding laundry, or even helping a person. Current Robots (VLAs): These robots are like a student who learns by rote memorization. You show them exactly how to perform a task, step-by-step, many times over. If you then ask them to do a slightly different version of the task, or in a new environment, they might get stuck because they've only memorized specific sequences of actions. They know what to do based on language, but not deeply how the physical world responds to their actions. DreamZero (WAM): DreamZero is like a student who learns by not just observing actions, but by also imagining the future. It watches tons of videos from the internet and robot demonstrations, learning not just the robot's movements, but also how the entire scene changes in response to those movements. It builds a "world model" in its head – a sense of physics, object interactions, and how things evolve. When you give DreamZero a new instruction (even for a task it's…

Continue to interactive post