机器人RSS 20252025

RT-2: Vision-Language-Action Models

Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David et al.Google DeepMind

摘要

We present RT-2, which transfers web-scale vision-language knowledge into robotic control. RT-2 directly translates visual and language inputs into robot actions, demonstrating emergent capabilities in object manipulation, semantic understanding, and cross-embodiment generalization.

vision-language-actionweb-scale knowledgerobotic controlemergent capabilitiessemantic manipulation

技术细节

仿真平台
测试机器人
FR3WidowX-250Saywer
模型骨架

PaLI-X / PaLM-E + Action Tokenizer

编码器

PaLI-X ViT-e + Language Encoder

解码器

Action Token Decoder (55B params, 32 layers)

相关公司(3)

京ICP备2026064258号-1