VLA ModelsarXiv preprint2025
AR-VLA: Autoregressive Action Expert for Vision-Language-Action Models
Darshan Nagendra Prasad, Lars Ullrich, Knut Graichen et al.Stanford University
摘要
We introduce AR-VLA, an autoregressive action expert designed for Vision-Language-Action models. Our approach uses autoregressive token prediction for continuous robot actions, enabling more flexible and compositional action generation compared to diffusion-based alternatives while maintaining strong manipulation performance.
autoregressive actionVLA modelvision-language-actionrobot manipulationtransformer
