WeTalkRobots · About

SteadyTray: Learning Object Balancing Tasks in Humanoid Tray Transport via Residual Reinforcement Learning

Humanoid Robots · 11/03/2026

Technical Analysis Stabilizing unsecured payloads against the inherent oscillations of dynamic bipedal locomotion remains a critical engineering bottleneck for humanoids in unstructured environments. To solve this, we introduce "ReST-RL", a hierarchical reinforcement learning architecture that explicitly decouples locomotion from payload stabilization High-level Analogy: Imagine you're a skilled waiter carrying a tray of filled glasses (the unsecured payload) through a busy restaurant (the robot's locomotion). If you try to consciously control every tiny muscle in your legs for walking and every tiny muscle in your arms to keep the tray perfectly level at the same time, it's incredibly difficult, and a sudden bump could cause a spill. ReST-RL is like having two specialized brains working together: Your 'Walking Brain' (the Base Policy): This brain is pre-trained and incredibly good at just walking, navigating, and keeping your body stable. It handles the core task of moving through the restaurant without falling. Your 'Tray-Stabilization Brain' (the Residual Module): This brain is only focused on the tray and glasses. It constantly monitors their balance. It doesn't tell your 'walking brain' how to walk, but rather sends small, subtle correction signals to your arms and torso. If a glass starts to tilt, this 'stabilization brain' makes a tiny, rapid adjustment to your arm's angle, without disrupting your smooth walking rhythm. Initially, the 'tray-stabilization brain' might…

Continue to interactive post