VLARSS 20252025
RoboVLM: Vision-Language Models for Robot Manipulation
Yonatan Bisk, Yuke Zhu, Dieter Fox, Animesh GargNVIDIA, UT Austin
摘要
We present RoboVLM, which adapts large vision-language models for direct robot manipulation control. RoboVLM fine-tunes pre-trained VLMs on robot demonstration data, enabling robots to understand complex visual scenes and follow detailed natural language manipulation instructions.
vision-language modelrobot manipulationfine-tuningvisual understandinglanguage instruction