AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Lianming, Hu, Haibo, Cui, Yufei, Zuo, Jiacheng, Wu, Shangyu, Guan, Nan, Xue, Chun Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909833564782592
author Huang, Lianming
Hu, Haibo
Cui, Yufei
Zuo, Jiacheng
Wu, Shangyu
Guan, Nan
Xue, Chun Jason
author_facet Huang, Lianming
Hu, Haibo
Cui, Yufei
Zuo, Jiacheng
Wu, Shangyu
Guan, Nan
Xue, Chun Jason
contents With the rapid advancement of autonomous driving, deploying Vision-Language Models (VLMs) to enhance perception and decision-making has become increasingly common. However, the real-time application of VLMs is hindered by high latency and computational overhead, limiting their effectiveness in time-critical driving scenarios. This challenge is particularly evident when VLMs exhibit over-inference, continuing to process unnecessary layers even after confident predictions have been reached. To address this inefficiency, we propose AD-EE, an Early Exit framework that incorporates domain characteristics of autonomous driving and leverages causal inference to identify optimal exit layers. We evaluate our method on large-scale real-world autonomous driving datasets, including Waymo and the corner-case-focused CODA, as well as on a real vehicle running the Autoware Universe platform. Extensive experiments across multiple VLMs show that our method significantly reduces latency, with maximum improvements reaching up to 57.58%, and enhances object detection accuracy, with maximum gains of up to 44%.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05404
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
Huang, Lianming
Hu, Haibo
Cui, Yufei
Zuo, Jiacheng
Wu, Shangyu
Guan, Nan
Xue, Chun Jason
Computer Vision and Pattern Recognition
Artificial Intelligence
With the rapid advancement of autonomous driving, deploying Vision-Language Models (VLMs) to enhance perception and decision-making has become increasingly common. However, the real-time application of VLMs is hindered by high latency and computational overhead, limiting their effectiveness in time-critical driving scenarios. This challenge is particularly evident when VLMs exhibit over-inference, continuing to process unnecessary layers even after confident predictions have been reached. To address this inefficiency, we propose AD-EE, an Early Exit framework that incorporates domain characteristics of autonomous driving and leverages causal inference to identify optimal exit layers. We evaluate our method on large-scale real-world autonomous driving datasets, including Waymo and the corner-case-focused CODA, as well as on a real vehicle running the Autoware Universe platform. Extensive experiments across multiple VLMs show that our method significantly reduces latency, with maximum improvements reaching up to 57.58%, and enhances object detection accuracy, with maximum gains of up to 44%.
title AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.05404