SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xuesong, Huang, Linjiang, Ma, Tao, Fang, Rongyao, Shi, Shaoshuai, Li, Hongsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913853278781440
author Chen, Xuesong
Huang, Linjiang
Ma, Tao
Fang, Rongyao
Shi, Shaoshuai
Li, Hongsheng
author_facet Chen, Xuesong
Huang, Linjiang
Ma, Tao
Fang, Rongyao
Shi, Shaoshuai
Li, Hongsheng
contents The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and common-sense reasoning. However, existing approaches often struggle with efficient integration and realtime decision-making due to computational demands. In this paper, we introduce SOLVE, an innovative framework that synergizes VLMs with end-to-end (E2E) models to enhance autonomous vehicle planning. Our approach emphasizes knowledge sharing at the feature level through a shared visual encoder, enabling comprehensive interaction between VLM and E2E components. We propose a Trajectory Chain-of-Thought (T-CoT) paradigm, which progressively refines trajectory predictions, reducing uncertainty and improving accuracy. By employing a temporal decoupling strategy, SOLVE achieves efficient cooperation by aligning high-quality VLM outputs with E2E real-time performance. Evaluated on the nuScenes dataset, our method demonstrates significant improvements in trajectory prediction accuracy, paving the way for more robust and reliable autonomous driving systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
Chen, Xuesong
Huang, Linjiang
Ma, Tao
Fang, Rongyao
Shi, Shaoshuai
Li, Hongsheng
Computer Vision and Pattern Recognition
The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and common-sense reasoning. However, existing approaches often struggle with efficient integration and realtime decision-making due to computational demands. In this paper, we introduce SOLVE, an innovative framework that synergizes VLMs with end-to-end (E2E) models to enhance autonomous vehicle planning. Our approach emphasizes knowledge sharing at the feature level through a shared visual encoder, enabling comprehensive interaction between VLM and E2E components. We propose a Trajectory Chain-of-Thought (T-CoT) paradigm, which progressively refines trajectory predictions, reducing uncertainty and improving accuracy. By employing a temporal decoupling strategy, SOLVE achieves efficient cooperation by aligning high-quality VLM outputs with E2E real-time performance. Evaluated on the nuScenes dataset, our method demonstrates significant improvements in trajectory prediction accuracy, paving the way for more robust and reliable autonomous driving systems.
title SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16805