Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xiong, Minhao, Wen, Zichen, Gu, Zhuangcheng, Liu, Xuyang, Zhang, Rui, Kang, Hengrui, Yang, Jiabing, Zhang, Junyuan, Li, Weijia, He, Conghui, Wang, Yafei, Zhang, Linfeng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912935905853440
author Xiong, Minhao
Wen, Zichen
Gu, Zhuangcheng
Liu, Xuyang
Zhang, Rui
Kang, Hengrui
Yang, Jiabing
Zhang, Junyuan
Li, Weijia
He, Conghui
Wang, Yafei
Zhang, Linfeng
author_facet Xiong, Minhao
Wen, Zichen
Gu, Zhuangcheng
Liu, Xuyang
Zhang, Rui
Kang, Hengrui
Yang, Jiabing
Zhang, Junyuan
Li, Weijia
He, Conghui
Wang, Yafei
Zhang, Linfeng
contents Vision-Language Models (VLMs) have emerged as a promising paradigm in autonomous driving (AD), providing a unified framework for perception and decision-making. However, their real-world deployment is hindered by significant computational overhead when processing high-resolution, multi-view images. This complexity stems from the massive number of visual tokens, which increases inference latency and memory consumption due to the quadratic complexity of self-attention. To address these challenges, we propose Prune2Drive, a plug-and-play visual token pruning framework for multi-view VLMs in AD. Prune2Drive introduces two core innovations: (i) a diversity-aware token selection mechanism that prioritizes semantic and spatial coverage across views, and (ii) a view-adaptive pruning controller that automatically learns optimal pruning ratios based on camera importance to downstream tasks. Unlike prior methods, Prune2Drive requires no model retraining or access to attention maps, ensuring compatibility with modern efficient attention implementations. Extensive experiments on the DriveLM and DriveLMM-o1 benchmarks demonstrate that Prune2Drive achieves significant speedups and memory savings with minimal performance impact. When retaining only 10% of visual tokens, our method achieves a 6.40x speedup in the prefilling phase and consumes only 13.4% of the original FLOPs, with a mere 3% average performance drop on the DriveLM benchmark. Code is available at: https://github.com/MinhaoXiong/Prune2Drive.git
format Preprint
id arxiv_https___arxiv_org_abs_2508_13305
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
Xiong, Minhao
Wen, Zichen
Gu, Zhuangcheng
Liu, Xuyang
Zhang, Rui
Kang, Hengrui
Yang, Jiabing
Zhang, Junyuan
Li, Weijia
He, Conghui
Wang, Yafei
Zhang, Linfeng
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) have emerged as a promising paradigm in autonomous driving (AD), providing a unified framework for perception and decision-making. However, their real-world deployment is hindered by significant computational overhead when processing high-resolution, multi-view images. This complexity stems from the massive number of visual tokens, which increases inference latency and memory consumption due to the quadratic complexity of self-attention. To address these challenges, we propose Prune2Drive, a plug-and-play visual token pruning framework for multi-view VLMs in AD. Prune2Drive introduces two core innovations: (i) a diversity-aware token selection mechanism that prioritizes semantic and spatial coverage across views, and (ii) a view-adaptive pruning controller that automatically learns optimal pruning ratios based on camera importance to downstream tasks. Unlike prior methods, Prune2Drive requires no model retraining or access to attention maps, ensuring compatibility with modern efficient attention implementations. Extensive experiments on the DriveLM and DriveLMM-o1 benchmarks demonstrate that Prune2Drive achieves significant speedups and memory savings with minimal performance impact. When retaining only 10% of visual tokens, our method achieves a 6.40x speedup in the prefilling phase and consumes only 13.4% of the original FLOPs, with a mere 3% average performance drop on the DriveLM benchmark. Code is available at: https://github.com/MinhaoXiong/Prune2Drive.git
title Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.13305