Technical Analysis is an environment-generalist visual loco-manipulation policy learned entirely in simulation and deployed zero-shot on real hardware, further extending it to multi-object training as an initial step toward a task generalist. High-level Analogy: Imagine you want to teach a robot to make a perfect omelet. Currently, robots can do simple tasks like stirring eggs or just walking, but combining complex walking (locomotion) and grabbing (manipulation) to make an omelet in any kitchen is extremely difficult. We usually rely on either endless real-world practice (expensive and messy) or perfectly scripted, step-by-step instructions in a virtual kitchen. FetchMan's Approach is like a student learning from a master chef in two stages: Stage 1: Learning from 'Cookbook Videos' (Behavior Cloning): The robot first watches countless videos of a master chef (a 'scripted controller') making omelets in a virtual kitchen. This chef always makes perfect omelets, but they have a secret internal checklist or 'phase switch' (e.g., 'Phase 1: Walk to Fridge', 'Phase 2: Grab Eggs', 'Phase 3: Cook'). The robot sees all the chef's movements, but it doesn't see the internal checklist. It tries its best to copy, and it gets pretty good, but it often struggles at the transitions – like trying to grab eggs before it's close enough to the fridge, or stirring an empty pan, because it doesn't know when the chef internally switches phases. It hits a performance ceiling, like a student who can…
FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences
Humanoid Robots · 29/08/2026