DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Anqing, Gao, Yu, Sun, Zhigang, Wang, Yiru, Wang, Jijun, Chai, Jinghao, Cao, Qian, Heng, Yuweng, Jiang, Hao, Dong, Yunda, Zhang, Zongzheng, Guo, Xianda, Sun, Hao, Zhao, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916773957206016
author Jiang, Anqing
Gao, Yu
Sun, Zhigang
Wang, Yiru
Wang, Jijun
Chai, Jinghao
Cao, Qian
Heng, Yuweng
Jiang, Hao
Dong, Yunda
Zhang, Zongzheng
Guo, Xianda
Sun, Hao
Zhao, Hao
author_facet Jiang, Anqing
Gao, Yu
Sun, Zhigang
Wang, Yiru
Wang, Jijun
Chai, Jinghao
Cao, Qian
Heng, Yuweng
Jiang, Hao
Dong, Yunda
Zhang, Zongzheng
Guo, Xianda
Sun, Hao
Zhao, Hao
contents Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called Diff-VLA. We explore the sparse diffusion representation for efficient multi-modal driving behavior. Moreover, we rethink the effectiveness of VLM driving decision and improve the trajectory generation guidance through deep interaction across agent, map instances and VLM output. Our method shows superior performance in Autonomous Grand Challenge 2025 which contains challenging real and reactive synthetic scenarios. Our methods achieves 45.0 PDMS.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19381
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
Jiang, Anqing
Gao, Yu
Sun, Zhigang
Wang, Yiru
Wang, Jijun
Chai, Jinghao
Cao, Qian
Heng, Yuweng
Jiang, Hao
Dong, Yunda
Zhang, Zongzheng
Guo, Xianda
Sun, Hao
Zhao, Hao
Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called Diff-VLA. We explore the sparse diffusion representation for efficient multi-modal driving behavior. Moreover, we rethink the effectiveness of VLM driving decision and improve the trajectory generation guidance through deep interaction across agent, map instances and VLM output. Our method shows superior performance in Autonomous Grand Challenge 2025 which contains challenging real and reactive synthetic scenarios. Our methods achieves 45.0 PDMS.
title DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2505.19381