VLARSS 20252025
OpenVLA-2: Open-Source Vision-Language-Action
Darshan Nagendra Prasad, Lars Ullrich, Knut Graichen et al.Stanford
摘要
We present OpenVLA-2, the next generation of open-source vision-language-action models for robotic control. OpenVLA-2 features improved visual grounding, better action prediction, and support for multiple robot embodiments, making it a versatile foundation for robotic manipulation research.
open-source VLArobotic controlvisual groundingaction predictionmultiple embodiments
技术细节
模型骨架
SigLIP-So400m + Llama 2 7B
编码器
SigLIP So400m/14 + Proprioceptive MLP
解码器
Llama 2 7B Action Decoder (32 layers)


