EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Yuhao, Shan, Bin, Ye, Xin, Chen, Cheng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917312302415872
author Chen, Yuhao
Shan, Bin
Ye, Xin
Chen, Cheng
author_facet Chen, Yuhao
Shan, Bin
Ye, Xin
Chen, Cheng
contents Multimodal Large Language Models (MLLMs) have shown strong performance in vision-language tasks, but their inference efficiency is severely limited by the exponential growth of visual tokens in complex scenarios such as high-resolution images and videos. Existing visual token pruning methods mainly operate after visual encoding, overlooking the substantial computational cost incurred during the encoding stage. To address this issue, we propose EvoPrune, an early-stage visual token pruning method for MLLMs that performs pruning directly during visual encoding. Specifically, EvoPrune employs a layer-wise pruning strategy guided by token similarity, diversity, and attention-based importance to retain the most informative visual tokens at selected encoding layers. Extensive experiments on image and video benchmarks validate the effectiveness of EvoPrune. In particular, on the VideoMME dataset, EvoPrune achieves 2$\times$ inference speedup with less than 1% performance degradation, demonstrating its potential for latency-sensitive MLLM deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03681
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
Chen, Yuhao
Shan, Bin
Ye, Xin
Chen, Cheng
Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10; I.2.6
Multimodal Large Language Models (MLLMs) have shown strong performance in vision-language tasks, but their inference efficiency is severely limited by the exponential growth of visual tokens in complex scenarios such as high-resolution images and videos. Existing visual token pruning methods mainly operate after visual encoding, overlooking the substantial computational cost incurred during the encoding stage. To address this issue, we propose EvoPrune, an early-stage visual token pruning method for MLLMs that performs pruning directly during visual encoding. Specifically, EvoPrune employs a layer-wise pruning strategy guided by token similarity, diversity, and attention-based importance to retain the most informative visual tokens at selected encoding layers. Extensive experiments on image and video benchmarks validate the effectiveness of EvoPrune. In particular, on the VideoMME dataset, EvoPrune achieves 2$\times$ inference speedup with less than 1% performance degradation, demonstrating its potential for latency-sensitive MLLM deployment.
title EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10; I.2.6
url https://arxiv.org/abs/2603.03681