Technical Analysis "ThinkAct" is a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. High-level Analogy: Imagine you're trying to teach a new chef (a robot) to cook a complex meal (a long-horizon task). Current methods are like either showing the chef many videos and hoping they mimic perfectly, or giving them an incredibly detailed, static recipe for every step. Both struggle if the kitchen setup changes or if the meal has many unpredictable steps. ThinkAct is like having two specialized experts working together: The Master Chef (Reasoning MLLM): This chef is brilliant at planning and thinking. You tell them the overall goal: 'Make a strawberry shortcake.' Instead of immediately acting, they think aloud, breaking down the task into logical steps ('First, pick the strawberries, then prepare the cream, then assemble'). Crucially, their thinking is constantly guided by visual feedback directly related to the actions. You show them 'pictures' of what the strawberries should look like at the start and end of being picked, and even a 'visual path' of how the spoon should move when mixing. If their mental plan looks visually unrealistic or misses a key step, you correct their 'thinking' immediately. This helps them create a robust, visually-grounded 'mental blueprint' (the visual plan latent). The Kitchen Assistant (Action Model): This assistant is great at executing precise actions, but needs clear,…
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
Reasoning VLA · 18/09/2025