VLANeurIPS 20252025
VLA-Bench: Comprehensive Benchmark for Vision-Language-Action Models
Karl Pertsch, Moo Jin Kim, Chelsea Finn, Sergey LevineStanford University, UC Berkeley
摘要
We present VLA-Bench, a comprehensive benchmark suite for evaluating vision-language-action models. VLA-Bench provides standardized tasks, environments, and metrics across multiple robot platforms, enabling fair comparison of VLA approaches and identifying key challenges for future research.
VLA benchmarkstandardized evaluationmulti-platformfair comparisoncomprehensive metrics
技术细节
数据集
VLA-Bench (100 tasks)Open X-Embodiment
测试机器人
模型骨架
Benchmark Framework (multiple baselines)
编码器
ViT-B/L/H + Proprioceptive Encoders
解码器
Transformer Decoders (4-32 layers)