VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Rui, Hu, Haofeng, Gao, Zhenhai, Liu, Jiaqiao, Fei, Gao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918497722826752
author Zhao, Rui
Hu, Haofeng
Gao, Zhenhai
Liu, Jiaqiao
Fei, Gao
author_facet Zhao, Rui
Hu, Haofeng
Gao, Zhenhai
Liu, Jiaqiao
Fei, Gao
contents Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieval-Augmented Generation (RAG) offers a solution by accessing external expert priors, standard visual retrieval suffers from high latency and semantic ambiguity. To address these challenges, we propose \textbf{VLADriver-RAG}, a framework that grounds planning in explicit, structure-aware historical knowledge. Specifically, we abstract sensory inputs into spatiotemporal semantic graphs via a \textit{Visual-to-Scenario} mechanism, effectively filtering visual noise. To ensure retrieval relevance, we employ a \textit{Scenario-Aligned Embedding Model} that utilizes Graph-DTW metric alignment to prioritize intrinsic topological consistency over superficial visual similarity. These retrieved priors are then fused within a query-based VLA backbone to synthesize precise, disentangled trajectories. Extensive experiments on the Bench2Drive benchmark establish a new state-of-the-art, achieving a Driving Score of 89.12.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08133
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
Zhao, Rui
Hu, Haofeng
Gao, Zhenhai
Liu, Jiaqiao
Fei, Gao
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieval-Augmented Generation (RAG) offers a solution by accessing external expert priors, standard visual retrieval suffers from high latency and semantic ambiguity. To address these challenges, we propose \textbf{VLADriver-RAG}, a framework that grounds planning in explicit, structure-aware historical knowledge. Specifically, we abstract sensory inputs into spatiotemporal semantic graphs via a \textit{Visual-to-Scenario} mechanism, effectively filtering visual noise. To ensure retrieval relevance, we employ a \textit{Scenario-Aligned Embedding Model} that utilizes Graph-DTW metric alignment to prioritize intrinsic topological consistency over superficial visual similarity. These retrieved priors are then fused within a query-based VLA backbone to synthesize precise, disentangled trajectories. Extensive experiments on the Bench2Drive benchmark establish a new state-of-the-art, achieving a Driving Score of 89.12.
title VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.08133