机器人RSS 20252025
RT-2: Vision-Language-Action Models
Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David et al.Google DeepMind
摘要
We present RT-2, which transfers web-scale vision-language knowledge into robotic control. RT-2 directly translates visual and language inputs into robot actions, demonstrating emergent capabilities in object manipulation, semantic understanding, and cross-embodiment generalization.
vision-language-actionweb-scale knowledgerobotic controlemergent capabilitiessemantic manipulation
技术细节
测试机器人
模型骨架
PaLI-X / PaLM-E + Action Tokenizer
编码器
PaLI-X ViT-e + Language Encoder
解码器
Action Token Decoder (55B params, 32 layers)


