BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Monninger, Thomas, Xie, Shaoyuan, Chen, Qi Alfred, Ding, Sihao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912948921827328
author Monninger, Thomas
Xie, Shaoyuan
Chen, Qi Alfred
Ding, Sihao
author_facet Monninger, Thomas
Xie, Shaoyuan
Chen, Qi Alfred
Ding, Sihao
contents The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios. However, existing methods typically feed LLMs with tokens from multi-view and multi-frame images independently, leading to redundant computation and limited spatial consistency. This separation in visual processing hinders accurate 3D spatial reasoning and fails to maintain geometric coherence across views. On the other hand, Bird's-Eye View (BEV) representations learned from geometrically annotated tasks (e.g., object detection) provide spatial structure but lack the semantic richness of foundation vision encoders. To bridge this gap, we propose BEVLM, a framework that connects a spatially consistent and semantically distilled BEV representation with LLMs. Through extensive experiments, we show that BEVLM enables LLMs to reason more effectively in cross-view driving scenes, improving accuracy by 46%, by leveraging BEV features as unified inputs. Furthermore, by distilling semantic knowledge from LLMs into BEV representations, BEVLM significantly improves closed-loop end-to-end driving performance by 29% in safety-critical scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06576
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
Monninger, Thomas
Xie, Shaoyuan
Chen, Qi Alfred
Ding, Sihao
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios. However, existing methods typically feed LLMs with tokens from multi-view and multi-frame images independently, leading to redundant computation and limited spatial consistency. This separation in visual processing hinders accurate 3D spatial reasoning and fails to maintain geometric coherence across views. On the other hand, Bird's-Eye View (BEV) representations learned from geometrically annotated tasks (e.g., object detection) provide spatial structure but lack the semantic richness of foundation vision encoders. To bridge this gap, we propose BEVLM, a framework that connects a spatially consistent and semantically distilled BEV representation with LLMs. Through extensive experiments, we show that BEVLM enables LLMs to reason more effectively in cross-view driving scenes, improving accuracy by 46%, by leveraging BEV features as unified inputs. Furthermore, by distilling semantic knowledge from LLMs into BEV representations, BEVLM significantly improves closed-loop end-to-end driving performance by 29% in safety-critical scenarios.
title BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2603.06576