VLACVPR 20252025

VLA-Perception: Enhanced Perception for Vision-Language-Action

Zhenjia Xu, Cheng Chi, Shuran Song, Xiaolong WangColumbia University, UC San Diego

摘要

We present VLA-Perception, which enhances the perception capabilities of vision-language-action models. VLA-Perception integrates advanced 3D perception modules including depth estimation, segmentation, and pose estimation, providing richer scene understanding for more accurate manipulation.

enhanced perceptionVLA3D perceptiondepth estimationscene understanding

技术细节

数据集
ScanNet++NYU Depth V2Ego4D
仿真平台
测试机器人
Stretch RE1Realsense-equipped Franka
模型骨架

Perception-Enhanced VLA

编码器

ViT-H/14 + Depth + Segmentation Encoder

解码器

Perception-Grounded Action Decoder (8 layers)

相关公司(3)

京ICP备2026064258号-1