VLARSS 20252025

RoboVLM: Vision-Language Models for Robot Manipulation

Yonatan Bisk, Yuke Zhu, Dieter Fox, Animesh GargNVIDIA, UT Austin

摘要

We present RoboVLM, which adapts large vision-language models for direct robot manipulation control. RoboVLM fine-tunes pre-trained VLMs on robot demonstration data, enabling robots to understand complex visual scenes and follow detailed natural language manipulation instructions.

vision-language modelrobot manipulationfine-tuningvisual understandinglanguage instruction

技术细节

仿真平台
测试机器人
FR3Kuka iiwaWidowX-250
模型骨架

PaLI-X + Action Head

编码器

PaLI-X ViT-e + Language Encoder

解码器

Action Token Decoder (12 layers)

京ICP备2026064258号-1