WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Songyan, Huang, Wenhui, Gao, Zihui, Chen, Hao, Lv, Chen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913615142977536
author Zhang, Songyan
Huang, Wenhui
Gao, Zihui
Chen, Hao
Lv, Chen
author_facet Zhang, Songyan
Huang, Wenhui
Gao, Zihui
Chen, Hao
Lv, Chen
contents The emergence of general human knowledge and impressive logical reasoning capacity in rapidly progressed vision-language models (VLMs) have driven increasing interest in applying VLMs to high-level autonomous driving tasks, such as scene understanding and decision-making. However, an in-depth study on the relationship between knowledge proficiency, especially essential driving expertise, and closed-loop autonomous driving performance requires further exploration. In this paper, we investigate the effects of the depth and breadth of fundamental driving knowledge on closed-loop trajectory planning and introduce WiseAD, a specialized VLM tailored for end-to-end autonomous driving capable of driving reasoning, action justification, object recognition, risk analysis, driving suggestions, and trajectory planning across diverse scenarios. We employ joint training on driving knowledge and planning datasets, enabling the model to perform knowledge-aligned trajectory planning accordingly. Extensive experiments indicate that as the diversity of driving knowledge extends, critical accidents are notably reduced, contributing 11.9% and 12.4% improvements in the driving score and route completion on the Carla closed-loop evaluations, achieving state-of-the-art performance. Moreover, WiseAD also demonstrates remarkable performance in knowledge evaluations on both in-domain and out-of-domain datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09951
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
Zhang, Songyan
Huang, Wenhui
Gao, Zihui
Chen, Hao
Lv, Chen
Computer Vision and Pattern Recognition
The emergence of general human knowledge and impressive logical reasoning capacity in rapidly progressed vision-language models (VLMs) have driven increasing interest in applying VLMs to high-level autonomous driving tasks, such as scene understanding and decision-making. However, an in-depth study on the relationship between knowledge proficiency, especially essential driving expertise, and closed-loop autonomous driving performance requires further exploration. In this paper, we investigate the effects of the depth and breadth of fundamental driving knowledge on closed-loop trajectory planning and introduce WiseAD, a specialized VLM tailored for end-to-end autonomous driving capable of driving reasoning, action justification, object recognition, risk analysis, driving suggestions, and trajectory planning across diverse scenarios. We employ joint training on driving knowledge and planning datasets, enabling the model to perform knowledge-aligned trajectory planning accordingly. Extensive experiments indicate that as the diversity of driving knowledge extends, critical accidents are notably reduced, contributing 11.9% and 12.4% improvements in the driving score and route completion on the Carla closed-loop evaluations, achieving state-of-the-art performance. Moreover, WiseAD also demonstrates remarkable performance in knowledge evaluations on both in-domain and out-of-domain datasets.
title WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09951