机器人arXiv2025

AtomVLA: Atomic Action Primitives for Vision-Language-Action Models

AtomVLA Team

摘要

We present AtomVLA, a framework that introduces atomic action primitives to enhance the reliability and interpretability of Vision-Language-Action (VLA) models. By decomposing complex manipulation tasks into a sequence of well-defined atomic actions, AtomVLA improves the compositional generalization of VLA models and provides clearer reasoning chains for robotic manipulation tasks.

atomic action primitivesVLA modelvision-language-actioncompositional generalizationinterpretabilityrobot manipulation
京ICP备2026064258号-1