Technical Analysis "Multi-Scale Embodied Memory" is an approach for mixed-modal long-horizon memory in robot policies. MEM combines video-based short-horizon memory, compressed via a video encoder, with text-based long-horizon memory. High-level Analogy: Imagine a super-smart robot chef trying to cook a complicated, multi-course meal. The 'Language Memory' is like the chef's well-organized recipe binder or a mental checklist. It doesn't remember every single tiny detail (like the exact angle of the spoon when adding salt). Instead, it keeps track of the big picture: 'I've added the potatoes,' 'The chicken is in the oven,' 'I've cleaned the first counter.' If the chef tries to add an ingredient but drops it, the recipe binder doesn't record the failure; it only gets updated when the ingredient is successfully added. This helps keep the 'recipe' short and focused on what's important for the long run. The 'Video Memory' is like the chef's ultra-short-term, detailed visual recall or a high-speed video replay in their mind. If they're trying to pick up a tricky ingredient and their hand blocks the view, or if they accidentally drop something, this 'video memory' instantly replays the last few seconds. This allows them to immediately adjust their grip, hand position, or movement to correct the mistake without having to consult the full recipe book. It's fast, visual, and helps with immediate, fine-grained actions and problem-solving. Together, these two memories (the recipe binder…
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Memory for VLA · 03/03/2026