Industry Brief 's Direct Video-Action Model (DVA) reformulates robot policies as video generation, unlocking data-efficient task learning, scaling, long-context memory, and one-shot learning. The News Rhoda AI has announced new research exploring the scalability of their Direct Video-to-Action (DVA) model, focusing on how model size and pre-training compute for video-based robot policies impact performance on real industrial manipulation tasks. What It Is The Direct Video-to-Action (DVA) model is a foundational robotics intelligence system developed by Rhoda AI. It redefines robot policies as video generation, utilizing hundreds of millions of openly accessible online videos for pre-training, similar to how large language models are trained on vast text data. This approach allows the robot to generate short video segments that predict its next movements, which are then converted into immediate actions (motor torques, joint angles). This closed-loop system continually generates and executes actions based on observed scenes, reacting quickly to real-world variations by making predictions in milliseconds. Built on this DVA architecture, Rhoda AI also launched FutureVision, a robotic intelligence platform designed to enable robots to operate reliably in dynamic real-world production environments where materials, layouts, and workflows are constantly changing. Unlike traditional vision-language-action (VLA) models, which often struggle with real-world variability, DVA learns a…
Rhoda AI Advances Robotics with 'Direct Video-to-Action' Model Trained on Web-Scale Video
From the industry · 10/03/2026