VLAICRA 20252025

MobileVLA: Vision-Language-Action for Mobile Manipulation

Tony Zhao, Zipeng Fu, Chelsea Finn, Jitendra MalikStanford University, UC Berkeley

摘要

We present MobileVLA, a vision-language-action model designed for mobile manipulation robots. MobileVLA jointly reasons about navigation and manipulation, enabling mobile robots to approach objects, reposition themselves, and perform manipulation tasks guided by natural language instructions.

mobile manipulationVLAnavigationjoint reasoninglanguage-guided

技术细节

数据集
SayCan DatasetOpen X-EmbodimentMobile Manipulation Data
测试机器人
模型骨架

Mobile-Optimized VLA (3B)

编码器

EfficientViT + Depth Encoder

解码器

Lightweight Action Decoder (4 layers)

相关公司(6)

京ICP备2026064258号-1