VLANeurIPS 20252025年10月

VLA-3D: Vision-Language-Action for 3D Manipulation

Yue Yang, Christopher Xie, Jiajun Wu, Fei-Fei LiStanford University

摘要

We present VLA-3D, a vision-language-action model that operates in 3D space for precise manipulation. VLA-3D processes 3D point clouds and language instructions to generate 6-DoF manipulation actions, achieving superior performance on precision assembly and insertion tasks compared to 2D-based approaches.

3D manipulationVLApoint cloud6-DoF actionprecision assembly

技术细节

数据集
Scan2CADPartNet-Mobility3D-FUTURE
模型骨架

3D-VLA + Point Cloud Backbone

编码器

PointNet++ + ViT-L/14 Multi-View

解码器

3D-Aware Action Decoder (8 layers)

京ICP备2026064258号-1