VLANeurIPS 20252025年10月
VLA-3D: Vision-Language-Action for 3D Manipulation
Yue Yang, Christopher Xie, Jiajun Wu, Fei-Fei LiStanford University
摘要
We present VLA-3D, a vision-language-action model that operates in 3D space for precise manipulation. VLA-3D processes 3D point clouds and language instructions to generate 6-DoF manipulation actions, achieving superior performance on precision assembly and insertion tasks compared to 2D-based approaches.
3D manipulationVLApoint cloud6-DoF actionprecision assembly
技术细节
数据集
Scan2CADPartNet-Mobility3D-FUTURE
模型骨架
3D-VLA + Point Cloud Backbone
编码器
PointNet++ + ViT-L/14 Multi-View
解码器
3D-Aware Action Decoder (8 layers)