WeTalkRobots · About

ROBOMETER: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

RL for VLA · 02/03/2026

Technical Analysis "ROBOMETER" is a scalable reward modeling framework that combines intra-trajectory progress supervision with inter-trajectory preference supervision. High-level Analogy: Imagine you're learning to become a judge for a talent show, like baking. Current robot reward models are like being shown only perfect cakes, and your job is to give a score from 0 to 1 for how 'perfect' each cake is. This works fine for perfect cakes, but what about cakes that are burnt, fallen, or just a bit lopsided? It's really hard to assign a precise 0-1 'perfection' score to every tiny mistake, and you learn very little from all those 'failed' attempts. ROBOMETER is like learning to judge cakes by not only seeing perfect examples but also by watching hundreds of bakers, some making perfect cakes, some making mistakes, and some outright failing. Instead of just scoring 'perfection', you also get to answer questions like: 'Which of these two cakes (one slightly burnt, one slightly lopsided) is better?' or 'Did this baker make more progress than that one, even if both failed?' By constantly comparing different attempts, even the bad ones, you develop a much deeper and more nuanced understanding of what makes a cake 'good' or 'bad', and you can even pinpoint where things started to go wrong. This way, you learn to give helpful feedback to bakers at all skill levels, not just the top experts. Motivation of the Work Existing general-purpose robot reward models typically learn to measure…

Continue to interactive post