DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yu, Jiang, Anqing, Wang, Yiru, Jijun, Wang, Jiang, Hao, Sun, Zhigang, Yuwen, Heng, Shuo, Wang, Zhao, Hao, Hao, Sun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908625803411456
author Gao, Yu
Jiang, Anqing
Wang, Yiru
Jijun, Wang
Jiang, Hao
Sun, Zhigang
Yuwen, Heng
Shuo, Wang
Zhao, Hao
Hao, Sun
author_facet Gao, Yu
Jiang, Anqing
Wang, Yiru
Jijun, Wang
Jiang, Hao
Sun, Zhigang
Yuwen, Heng
Shuo, Wang
Zhao, Hao
Hao, Sun
contents Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of essential world knowledge to understand and reason about surrounding environments. In contrast, Vision-Language-Action (VLA) models leverage world knowledge to handle challenging cases, but their limited 3D reasoning capability can lead to physically infeasible actions. In this work we introduce DiffVLA++, an enhanced autonomous driving framework that explicitly bridges cognitive reasoning and E2E planning through metric-guided alignment. First, we build a VLA module directly generating semantically grounded driving trajectories. Second, we design an E2E module with a dense trajectory vocabulary that ensures physical feasibility. Third, and most critically, we introduce a metric-guided trajectory scorer that guides and aligns the outputs of the VLA and E2E modules, thereby integrating their complementary strengths. The experiment on the ICCV 2025 Autonomous Grand Challenge leaderboard shows that DiffVLA++ achieves EPDMS of 49.12.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17148
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
Gao, Yu
Jiang, Anqing
Wang, Yiru
Jijun, Wang
Jiang, Hao
Sun, Zhigang
Yuwen, Heng
Shuo, Wang
Zhao, Hao
Hao, Sun
Robotics
Computer Vision and Pattern Recognition
Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of essential world knowledge to understand and reason about surrounding environments. In contrast, Vision-Language-Action (VLA) models leverage world knowledge to handle challenging cases, but their limited 3D reasoning capability can lead to physically infeasible actions. In this work we introduce DiffVLA++, an enhanced autonomous driving framework that explicitly bridges cognitive reasoning and E2E planning through metric-guided alignment. First, we build a VLA module directly generating semantically grounded driving trajectories. Second, we design an E2E module with a dense trajectory vocabulary that ensures physical feasibility. Third, and most critically, we introduce a metric-guided trajectory scorer that guides and aligns the outputs of the VLA and E2E modules, thereby integrating their complementary strengths. The experiment on the ICCV 2025 Autonomous Grand Challenge leaderboard shows that DiffVLA++ achieves EPDMS of 49.12.
title DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.17148