VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guo, Ziang, Gubernatorov, Konstantin, Asfaw, Selamawit, Yagudin, Zakhar, Tsetserukou, Dzmitry
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929737852518400
author Guo, Ziang
Gubernatorov, Konstantin
Asfaw, Selamawit
Yagudin, Zakhar
Tsetserukou, Dzmitry
author_facet Guo, Ziang
Gubernatorov, Konstantin
Asfaw, Selamawit
Yagudin, Zakhar
Tsetserukou, Dzmitry
contents In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing with the representation of state-action mapping in the end-to-end autonomous driving paradigm, we introduce a novel pipeline, VDT-Auto. Leveraging the advancement of the state understanding of Visual Language Model (VLM), incorporating with diffusion Transformer-based action generation, our VDT-Auto parses the environment geometrically and contextually for the conditioning of the diffusion process. Geometrically, we use a bird's-eye view (BEV) encoder to extract feature grids from the surrounding images. Contextually, the structured output of our fine-tuned VLM is processed into textual embeddings and noisy paths. During our diffusion process, the added noise for the forward process is sampled from the noisy path output of the fine-tuned VLM, while the extracted BEV feature grids and embedded texts condition the reverse process of our diffusion Transformers. Our VDT-Auto achieved 0.52m on average L2 errors and 21% on average collision rate in the nuScenes open-loop planning evaluation. Moreover, the real-world demonstration exhibited prominent generalizability of our VDT-Auto. The code and dataset will be released after acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
Guo, Ziang
Gubernatorov, Konstantin
Asfaw, Selamawit
Yagudin, Zakhar
Tsetserukou, Dzmitry
Computer Vision and Pattern Recognition
Robotics
In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing with the representation of state-action mapping in the end-to-end autonomous driving paradigm, we introduce a novel pipeline, VDT-Auto. Leveraging the advancement of the state understanding of Visual Language Model (VLM), incorporating with diffusion Transformer-based action generation, our VDT-Auto parses the environment geometrically and contextually for the conditioning of the diffusion process. Geometrically, we use a bird's-eye view (BEV) encoder to extract feature grids from the surrounding images. Contextually, the structured output of our fine-tuned VLM is processed into textual embeddings and noisy paths. During our diffusion process, the added noise for the forward process is sampled from the noisy path output of the fine-tuned VLM, while the extracted BEV feature grids and embedded texts condition the reverse process of our diffusion Transformers. Our VDT-Auto achieved 0.52m on average L2 errors and 21% on average collision rate in the nuScenes open-loop planning evaluation. Moreover, the real-world demonstration exhibited prominent generalizability of our VDT-Auto. The code and dataset will be released after acceptance.
title VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2502.20108