VLACVPR 20252025
VLA-Perception: Enhanced Perception for Vision-Language-Action
Zhenjia Xu, Cheng Chi, Shuran Song, Xiaolong WangColumbia University, UC San Diego
摘要
We present VLA-Perception, which enhances the perception capabilities of vision-language-action models. VLA-Perception integrates advanced 3D perception modules including depth estimation, segmentation, and pose estimation, providing richer scene understanding for more accurate manipulation.
enhanced perceptionVLA3D perceptiondepth estimationscene understanding
技术细节
数据集
ScanNet++NYU Depth V2Ego4D
测试机器人
Stretch RE1Realsense-equipped Franka
模型骨架
Perception-Enhanced VLA
编码器
ViT-H/14 + Depth + Segmentation Encoder
解码器
Perception-Grounded Action Decoder (8 layers)


