WeTalkRobots · About

Ψ0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

VLA Loco-manipulation · 12/03/2026

Technical Analysis "Psi-Zero" is an open foundation model to address challenging humanoid loco-manipulation tasks. First, autoregressively pretrain a VLM backbone on large-scale egocentric human videos to acquire generalizable visual-action representations. Then, post-train a flow-based action expert on high-quality humanoid robot data to learn precise robot joint control. High-level Analogy: Imagine you want to teach a robot how to perform complex tasks, like cooking in a kitchen. The old way was like showing the robot a mix of cooking videos filmed by people, and then also trying to directly move the robot's arms and legs at the same time. The problem? Human hands and robot hands are built differently, and the robot gets confused trying to learn both 'what to do' and 'how to move its unique body' all at once from jumbled data. It ends up being slow and not very good. Ψ0's new approach is smarter: First, learn from watching (VLM Pre-training): The robot first watches hundreds of hours of high-quality cooking videos from a human's point of view – like watching a master chef's GoPro footage. It doesn't try to move yet, but it learns the general steps, intentions, and visual cues for cooking tasks (e.g., 'grasp the knife,' 'stir the pot'). It's like learning the 'language' and 'flow' of cooking without doing it. Then, learn precise movements (Action Expert Post-training): Once it understands the concepts from watching, it's given a smaller amount of very high-quality practice…

Continue to interactive post