SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Peizheng, Zhang, Zhenghao, Holtz, David, Yu, Hang, Yang, Yutong, Lai, Yuzhi, Song, Rui, Geiger, Andreas, Zell, Andreas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913150537826304
author Li, Peizheng
Zhang, Zhenghao
Holtz, David
Yu, Hang
Yang, Yutong
Lai, Yuzhi
Song, Rui
Geiger, Andreas
Zell, Andreas
author_facet Li, Peizheng
Zhang, Zhenghao
Holtz, David
Yu, Hang
Yang, Yutong
Lai, Yuzhi
Song, Rui
Geiger, Andreas
Zell, Andreas
contents End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-grained 3D spatial relationships which is a fundamental requirement for systems interacting with the physical world. To address this issue, we propose SpaceDrive, a spatial-aware VLM-based driving framework that treats spatial information as explicit positional encodings (PEs) instead of textual digit tokens, enabling joint reasoning over semantic and spatial representations. SpaceDrive employs a universal positional encoder to all 3D coordinates derived from multi-view depth estimation, historical ego-states, and text prompts. These 3D PEs are first superimposed to augment the corresponding 2D visual tokens. Meanwhile, they serve as a task-agnostic coordinate representation, replacing the digit-wise numerical tokens as both inputs and outputs for the VLM. This mechanism enables the model to better index specific visual semantics in spatial reasoning and directly regress trajectory coordinates rather than generating digit-by-digit, thereby enhancing planning accuracy. Extensive experiments validate that SpaceDrive achieves state-of-the-art open-loop performance on the nuScenes dataset and the second-best Driving Score of 78.02 on the Bench2Drive closed-loop benchmark over existing VLM-based methods. Code is available at: https://github.com/zhenghao2519/SpaceDrive.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10719
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
Li, Peizheng
Zhang, Zhenghao
Holtz, David
Yu, Hang
Yang, Yutong
Lai, Yuzhi
Song, Rui
Geiger, Andreas
Zell, Andreas
Computer Vision and Pattern Recognition
End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-grained 3D spatial relationships which is a fundamental requirement for systems interacting with the physical world. To address this issue, we propose SpaceDrive, a spatial-aware VLM-based driving framework that treats spatial information as explicit positional encodings (PEs) instead of textual digit tokens, enabling joint reasoning over semantic and spatial representations. SpaceDrive employs a universal positional encoder to all 3D coordinates derived from multi-view depth estimation, historical ego-states, and text prompts. These 3D PEs are first superimposed to augment the corresponding 2D visual tokens. Meanwhile, they serve as a task-agnostic coordinate representation, replacing the digit-wise numerical tokens as both inputs and outputs for the VLM. This mechanism enables the model to better index specific visual semantics in spatial reasoning and directly regress trajectory coordinates rather than generating digit-by-digit, thereby enhancing planning accuracy. Extensive experiments validate that SpaceDrive achieves state-of-the-art open-loop performance on the nuScenes dataset and the second-best Driving Score of 78.02 on the Bench2Drive closed-loop benchmark over existing VLM-based methods. Code is available at: https://github.com/zhenghao2519/SpaceDrive.
title SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10719