WeTalkRobots · About

QWEN Embodied AI Model Family

Reports · 01/09/2026

Qwen-RobotSuite: A Leap Towards Embodied AI The Qwen-RobotSuite is a Vision-Language-Action (VLA) foundation model specifically designed to tackle complex robotic manipulation tasks. It enables robots to handle objects and perform precise actions using various robot arms and grippers, aiming for strong generalization across different robot embodiments and task types. Data Used The model was trained on an extensive corpus of approximately 38,100 hours of manipulation data. This includes a combination of open-source robotic datasets and human demonstration videos. To address the scarcity and heterogeneity of robot data, Qwen-RobotManip utilizes a human-to-robot synthesis pipeline. This pipeline converts egocentric human manipulation videos into robot demonstrations, allowing for training across at least 15 different robot platforms. Architecture Qwen-RobotManip is built upon a Qwen3.5-4B (or Qwen-VL) vision-language backbone. Its architecture features a unified alignment framework that standardizes representations, motions, and behaviors across diverse manipulation data. The model couples this VLM backbone with a flow-matching Diffusion Transformer (DiT) action head. Actions for end-effectors are represented as camera-frame delta poses, and a unified 80-dimensional canonical representation is used for both states and actions. Qwen-RobotNav: Intelligent Navigation for Agentic Systems Qwen-RobotNav is a scalable navigation model designed for agentic systems, addressing a wide…

Continue to interactive post