无人驾驶CVPR 20252025

DriveVLM: Vision-Language Model for Autonomous Driving

Yupeng Zheng, Yilun Chen, Zhaoxiang Zhang, Jiwen LuTsinghua University, Chinese Academy of Sciences

摘要

We present DriveVLM, which applies vision-language models to autonomous driving decision-making. DriveVLM reasons about driving scenes using natural language, enabling interpretable decision-making that can explain its actions and handle rare or novel driving scenarios through language-based reasoning.

vision-language modeldriving decisioninterpretable AIlanguage reasoningnovel scenario

技术细节

模型骨架

VLM (LLaVA-13B) + BEV Encoder

编码器

CLIP ViT-L/14 + BEVFormer Encoder

解码器

LLaVA-13B Language Decoder + Action Head (40+4 layers)

相关公司(1)

京ICP备2026064258号-1