<h1 id="technical" Technical Analysis</h1 Jump to Section: Motivation motivation | Results summary-of-results | Conclusions final-conclusions "DiT4DiT" https://dit4dit.github.io/ is an end-to-end Video- Action Model that couples a video Diffusion Transformer with an action Diffusion Transformer in a unified cascaded framework. Instead of relying on reconstructed future frames, DiT4DiT extracts intermediate denoising features from the video generation process and uses them as temporally grounded conditions for action prediction. High-level Analogy : Imagine teaching a robot to cook. <br <b Traditional methods VLA </b are like giving the robot a cookbook text instructions and pictures of ingredients static images . It knows what to do and what things look like, but not intuitively how they behave or how to physically interact with them e.g., how to stir dough, how an egg breaks . It has to learn all the real-world physics from scratch through trial and error. <br This new approach <b DiT4DiT</b is like giving the robot two specialized chefs who work together: <br 1. <b The 'Visual Planner' Video DiT </b : This chef specializes in imagining future cooking steps. If you tell it to 'chop vegetables,' it doesn't just show you a final picture of chopped veggies. It mentally simulates the entire process : the knife moving, the vegetable cutting, how the pieces will fall. <br 2. <b The 'Action Executor' Action DiT </b : This chef's job is to control the robot's hands. Crucially, it do
DIT4DIT: JOINTLY MODELING VIDEO DYNAMICS AND ACTIONS FOR GENERALIZABLE ROBOT CONTROL
World Action Models · 22/03/2026