Scholay

学术搜索 · AI 审稿 · LaTeX 协作

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

作者:Felix Brandstaetter, Erik Schütz, Katharina Winter, Fabian B. Flohr · 发表于:2025 IEEE Intelligent Vehicles Symposium (IV) · 年份:2025 · DOI:10.1109/iv64158.2025.11097781 · 被引用次数:8 · 研究领域:Computer Science

Autonomous driving technology has the potential to transform transportation, but its wide adoption depends on the development of interpretable and transparent decision-making systems. Scene captioning, which generates natural language descriptions of the driving environment, plays a crucial role in enhancing transparency, safety, and human-AI interaction. We introduce BEV-LLM, a lightweight model for 3D captioning of autonomous driving scenes. BEV-LLM leverages BEVFusion to combine 3D LiDAR point clouds and multi-view images, incorporating a novel absolute positional encoding for view-specific scene descriptions. Despite using a small 1B parameter base model, BEV-LLM achieves competitive performance on the nuCaption dataset, surpassing state-of-the-art by up to 5% in BLEU scores. Additionally, we release two new datasets — nu-View (focused on environmental conditions and viewpoints) and GroundView (focused on object grounding) — to better assess scene captioning across diverse driving scenarios and address gaps in current benchmarks, along with initial benchmarking results demonstrating their effectiveness.