无人驾驶CVPR 20252025
DriveVLM: Vision-Language Model for Autonomous Driving
Yupeng Zheng, Yilun Chen, Zhaoxiang Zhang, Jiwen LuTsinghua University, Chinese Academy of Sciences
摘要
We present DriveVLM, which applies vision-language models to autonomous driving decision-making. DriveVLM reasons about driving scenes using natural language, enabling interpretable decision-making that can explain its actions and handle rare or novel driving scenarios through language-based reasoning.
vision-language modeldriving decisioninterpretable AIlanguage reasoningnovel scenario
技术细节
数据集
模型骨架
VLM (LLaVA-13B) + BEV Encoder
编码器
CLIP ViT-L/14 + BEVFormer Encoder
解码器
LLaVA-13B Language Decoder + Action Head (40+4 layers)
