WeTalkRobots · About

RDT-1: Robotic Diffusion Transformer

VLA · 0

Technical Analysis RDT-1 demo video The paper introduces the Robotics Diffusion Transformer (RDT), a diffusion foundation model for bimanual manipulation. The authors address the key challenges of data scarcity and the complexity of coordinating two robot arms, which leads to multi-modal action distributions. Motivation of the Work Current methods for bimanual manipulation are often limited by a lack of data and struggle with the large, multi-modal action space of two-armed robots. While foundation models have shown promise, their development for bimanual tasks has been hindered by the high cost of dual-arm systems, leading to severe data scarcity. The goal of this work is to overcome these challenges and create a foundation model that can generalize across different tasks, objects, and environments. Comparison to Other Works and the State of the Art Learning-based Bimanual Manipulation: Previous methods in this area often use simplified models or introduce biases to reduce the action space, which limits their versatility and ability to capture multi-modality. Foundation Models for Robotics: Most existing models adapt large vision-language models, which can lead to issues like quantization errors and uncoordinated behaviors. In contrast, RDT uses a diffusion model to handle continuous control and multi-modality more effectively. RDT is also significantly larger than previous models, scaled up to 1.2 billion parameters, compared to a Transformer-based diffusion policy with up…

Continue to interactive post