Vision-Language-Action (VLA) and World Action Models (WAMs): An Overview The field of robotics is rapidly advancing towards creating generalist agents capable of performing a wide array of tasks in diverse environments. Two prominent paradigms are driving this progress: Vision-Language-Action (VLA) models and World Action Models (WAMs). While both aim to bridge the gap between high-level human instructions and low-level robot control, they take fundamentally different approaches — especially in how they model and interact with physical reality. As NVIDIA's technical blog Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models puts it, WAMs are now "emerging as a second major recipe for robot foundation models alongside VLM-based VLA models" — a signal that the field is actively bifurcating around these two philosophies. --- Vision-Language-Action (VLA) Models: Bridging Perception and Action VLA models extend the capabilities of large Vision-Language Models (VLMs) to enable robots to perform physical actions. Their core idea is to leverage the vast semantic knowledge and reasoning abilities acquired by VLMs from massive internet-scale image and text data, and then adapt this understanding to generate robot-specific motor commands. Core Architecture and Functionality VLAs typically integrate a pre-trained VLM backbone with a specialized action-generation head. The VLM processes visual observations (camera images) and natural language instructions, producing…
VLAs and WAMs: The Brain Architectures of Embodied Ai
Reports · 01/10/2026