WeTalkRobots curates and summarizes the latest advancements in AI, robotics, humanoid control, reinforcement learning, and vision-language models. Explore research papers, podcasts, and videos focused on robot learning, manipulation, and autonomous systems. It holds a curated list of research papers, with summaries, images, and audio podcasts.
Browse some posts: Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization, BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion, EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data, EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos, F1VLA, GigaWorld-0: World Models as Data Engine to Empower Embodied AI, GigaWorld-Policy: An Efficient Action-Centered World–Action Model, InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions, Interactive World Simulator for Robot Policy Training and Evaluation, Masquerade: Learning from In-the-wild Human Videos using Data-Editing, OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control, Phantom: Training Robots Without Robots Using Only Human Videos, MEM: Multi-Scale Embodied Memory for Vision Language Action Models, π∗0.6: a VLA That Learns From Experience, Ψ0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation, RL Token: Bootstrapping Online RL with Vision-Language-Action Models, ROBOMETER: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons, SteadyTray: Learning Object Balancing Tasks in Humanoid Tray Transport via Residual Reinforcement Learning, ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning, TidyBot-Universe: A Modular and Service-Oriented Framework for Robotic Tidy-Up Tasks, TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System, Evaluating Gemini Robotics Policies in a Veo World Simulator, Vision-Language Agent (VLA) from Scratch, Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos, WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL, Scalable and General Whole-Body Control for Cross-Humanoid Locomotion, ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents, ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning, ASAP: Aligning Simulation and Real-world Physics, Being-H0.7: A Latent World-Action Model from Egocentric Videos, CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation, DeepMimic, Dexbotic: Open-Source Vision-Language-Action Toolbox, DIT4DIT: JOINTLY MODELING VIDEO DYNAMICS AND ACTIONS FOR GENERALIZABLE ROBOT CONTROL, DreamVLA, Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera, EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration, Embodied COT, Fast-WAM: Do World Action Models Need Test-time Future Imagination?, Figure Helix VLA, Reconstructing Hands in 3D with Transformers, HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos, Hitter: Humanoid Table Tennis Robot, VLAb: A Modular and Extensible Research Platform for Vision-Language Models, Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack, LARGE VIDEO PLANNER ENABLES GENERALIZABLE ROBOT CONTROL, LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction, mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs, NaVILA: Legged Robot Vision-Language-Action Model for Navigation, COMPASS: Cross-embOdiment Mobility Policy via ResiduAl RL and Skill Synthesis, DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos, DREAMGEN: Unlocking Generalization in Robot Learning through Video World Models, World Action Models are Zero-shot Policies, ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation, SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control, OmniRetarget, Open-source Robot VLAs and VLMs, π0.5: a Vision-Language-Action Model with Open-World Generalization, π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities, π0: A Vision-Language-Action Flow Model for General Robot Control, π3: PERMUTATION-EQUIVARIANT VISUAL GEOMETRY LEARNING, PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models, RDT-1: Robotic Diffusion Transformer, Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision–Language–Action Models via Latent Iterative Reasoning, ResMimic, RoboClaw: A Universal and Flexible Robotic Claw System for High-Performance Grasping, RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation, SoftMimic: Learning Compliant Whole-body Control from Examples, StarVLA: VLA Model Codebase, TWIST: Teleoperated Whole-Body Imitation System, UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers, VGGT-Ω, WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild, Cross-Hand Latent Representation for Vision-Language-Action Models, Igniting VLMs toward the Embodied Space, ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training, Flexion, Niantic Spatial, and NVIDIA Accelerate Humanoid Robot Deployment with Zero-Shot Sim-to-Real Transfer
Categories: Learning from Humans, Humanoid Robots, VLA, Video Generative Models, World Action Models, Memory for VLA, RL for VLA, VLA Loco-manipulation, Reasoning VLA, Agentic AI, Open-source VLA, Cross-embodiment Learning, 3D Reconstruction, Legged Robots, Navigation, From the industry