VLANeurIPS 20252025
ActionVLA: Action-Centric Vision-Language-Action Pre-training
Karl Pertsch, Moo Jin Kim, Chelsea Finn, Sergey LevineStanford University, UC Berkeley
摘要
We present ActionVLA, an action-centric pre-training approach for vision-language-action models. ActionVLA pre-trains on large-scale action-annotated video data, learning action representations that transfer effectively to diverse robot manipulation tasks with minimal fine-tuning data.
action-centricVLA pre-trainingaction representationvideo dataminimal fine-tuning
技术细节
测试机器人
模型骨架
Action Tokenizer + VLM Backbone
编码器
ViT-L/14 + Action Encoder
解码器
Action-Centric Transformer Decoder (8 layers)
