<h1 id="technical" Technical Analysis</h1 Jump to Section: Motivation motivation | Results summary-of-results | Conclusions final-conclusions <figure class="my-6" <video controls playsinline style="display: block; margin: 0 auto; max-width: 100%; height: auto;" src="$BUCKET URL/Pi-MEME/videos/video in post.mp4" </video </figure "Multi-Scale Embodied Memory" https://www.pi.website/research/memory is an approach for mixed-modal long-horizon memory in robot policies. MEM combines video-based short-horizon memory, compressed via a video encoder, with text-based long-horizon memory. High-level Analogy : Imagine a super-smart robot chef trying to cook a complicated, multi-course meal. The 'Language Memory' is like the chef's well-organized recipe binder or a mental checklist. It doesn't remember every single tiny detail like the exact angle of the spoon when adding salt . Instead, it keeps track of the big picture: 'I've added the potatoes,' 'The chicken is in the oven,' 'I've cleaned the first counter.' If the chef tries to add an ingredient but drops it, the recipe binder doesn't record the failure; it only gets updated when the ingredient is successfully added. This helps keep the 'recipe' short and focused on what's important for the long run. The 'Video Memory' is like the chef's ultra-short-term, detailed visual recall or a high-speed video replay in their mind. If they're trying to pick up a tricky ingredient and their hand blocks the view, or if they accidentally drop somet
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Memory for VLA · 03/03/2026