Technical Analysis "VITRA" is a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot end-effector, we show that “in the-wild” egocentric human videos without any annotations can be transformed into data formats fully aligned with existing robotic V-L-A training data in terms of task granularity and labels. High-level Analogy: Imagine you want to teach a highly advanced robot to perform countless everyday tasks with its hands, like opening a jar, stirring a pot, or picking up a small, unusually shaped toy. Instead of manually programming each precise movement or spending years painstakingly demonstrating every single task in a controlled lab setting, what if you could simply let the robot watch millions of home videos of humans naturally performing these actions in their daily lives? This paper is like building a smart system that watches all those diverse human videos, automatically understands what each hand is doing, and translates those movements into detailed instructions that a robot can understand. After watching and 'learning' from so much real-world human experience, the robot gets a general sense of how hands interact with objects. Then, with just a tiny bit of specific practice in its own robot body, it becomes incredibly good and adaptable at performing new, complex tasks, even with objects it has never…
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Learning from Humans · 24/10/2025