WeTalkRobots · About

MolmoAct

VLA · 05/05/2026

MolmoAct: From Action Reasoning to Real-World Deployment The Core Problem: Robots That Act Without Thinking Most robotics foundation models follow the same basic recipe: take an image, read an instruction, output a motor command. This direct perception-to-action mapping has driven real progress, but it has a fundamental limitation — the model offers no intermediate reasoning, no way to inspect why it chose one motion over another, and limited ability to generalize when the scene changes. At the same time, language models made a decisive shift away from brute-force scaling toward structured reasoning — chain-of-thought, intermediate representations, and step-by-step planning. MolmoAct asks: can robotics do the same? The answer, developed across two generations of models by researchers at the Allen Institute for AI (AI2) and the University of Washington, is yes — and the path from MolmoAct v1 to MolmoAct2 is a story of making that reasoning faster, more grounded, and deployable in the real world. --- MolmoAct v1: Reasoning in Space The Big Idea: Action Reasoning Models (ARMs) MolmoAct introduces the concept of an Action Reasoning Model (ARM) — a robotic foundation model that doesn't jump straight from pixels to actions. Instead, it works through a structured, three-stage reasoning chain: Depth Perception Tokens — sense and reconstruct the 3D structure of the scene Visual Reasoning Trace — sketch the intended end-effector path as a 2D trajectory overlay Action Tokens — predict…

Continue to interactive post