机器人ICLR 20252025
RT-3: Scaling Robot Foundation Models via Multitask Learning
Ted Xiao, Fei Xia, Jonathan Tompson, Brianna Zitkovich, Pierre SermanetGoogle DeepMind
摘要
We present RT-3, a scaled robot foundation model trained on diverse multitask data spanning manipulation, navigation, and locomotion. RT-3 demonstrates strong zero-shot generalization to novel objects and environments, achieving state-of-the-art performance across 50+ real-world robot tasks with a single unified policy.
robot foundation modelmultitask learningzero-shot generalizationmanipulationscaling
技术细节
数据集
模型骨架
PaLI-X 55B + Action Tokenizer
编码器
PaLI-X ViT-e + Language Encoder
解码器
Action Token Decoder (finetuned, 32 layers)