VLAICRA 20252025
MobileVLA: Vision-Language-Action for Mobile Manipulation
Tony Zhao, Zipeng Fu, Chelsea Finn, Jitendra MalikStanford University, UC Berkeley
摘要
We present MobileVLA, a vision-language-action model designed for mobile manipulation robots. MobileVLA jointly reasons about navigation and manipulation, enabling mobile robots to approach objects, reposition themselves, and perform manipulation tasks guided by natural language instructions.
mobile manipulationVLAnavigationjoint reasoninglanguage-guided
技术细节
数据集
测试机器人
模型骨架
Mobile-Optimized VLA (3B)
编码器
EfficientViT + Depth Encoder
解码器
Lightweight Action Decoder (4 layers)





