HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Qihui, Zhang, Tao, Wang, Yuchen, Wen, Zijian, Zhang, Mengjie, Chen, Shuangwu, Tan, Xiaobin, Yang, Jian, Liu, Yang, Dong, Zhenhua, Yu, Xianzhi, Pan, Yinfei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917394627166208
author Zhu, Qihui
Zhang, Tao
Wang, Yuchen
Wen, Zijian
Zhang, Mengjie
Chen, Shuangwu
Tan, Xiaobin
Yang, Jian
Liu, Yang
Dong, Zhenhua
Yu, Xianzhi
Pan, Yinfei
author_facet Zhu, Qihui
Zhang, Tao
Wang, Yuchen
Wen, Zijian
Zhang, Mengjie
Chen, Shuangwu
Tan, Xiaobin
Yang, Jian
Liu, Yang
Dong, Zhenhua
Yu, Xianzhi
Pan, Yinfei
contents In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time or resource-constrained applications. Visual token pruning is a promising strategy for reducing the cost of MLLM inference by removing redundant visual tokens. Existing research usually assumes that all attention heads contribute equally to the visual interpretation. However, our study reveals that different heads may capture distinct visual semantics and inherently play distinct roles in visual processing. In light of this observation, we propose HAWK, a head importance-aware visual token pruning method that perceives the varying importance of attention heads in visual tasks to maximize the retention of crucial tokens. By leveraging head importance weights and text-guided attention to assess visual token significance, HAWK effectively retains task-relevant visual tokens while removing redundant ones. The proposed HAWK is entirely training-free and can be seamlessly applied to various MLLMs. Extensive experiments on multiple mainstream vision-language benchmarks demonstrate that HAWK achieves state-of-the-art accuracy. When applied to Qwen2.5-VL, HAWK retains 96.0% of the original accuracy after pruning 80.2% of the visual tokens. Additionally, it reduces end-to-end latency to 74.4% of the original and further decreases GPU memory usage across the tested models. The code is available at https://github.com/peppery77/HAWK.git.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07812
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
Zhu, Qihui
Zhang, Tao
Wang, Yuchen
Wen, Zijian
Zhang, Mengjie
Chen, Shuangwu
Tan, Xiaobin
Yang, Jian
Liu, Yang
Dong, Zhenhua
Yu, Xianzhi
Pan, Yinfei
Computer Vision and Pattern Recognition
In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time or resource-constrained applications. Visual token pruning is a promising strategy for reducing the cost of MLLM inference by removing redundant visual tokens. Existing research usually assumes that all attention heads contribute equally to the visual interpretation. However, our study reveals that different heads may capture distinct visual semantics and inherently play distinct roles in visual processing. In light of this observation, we propose HAWK, a head importance-aware visual token pruning method that perceives the varying importance of attention heads in visual tasks to maximize the retention of crucial tokens. By leveraging head importance weights and text-guided attention to assess visual token significance, HAWK effectively retains task-relevant visual tokens while removing redundant ones. The proposed HAWK is entirely training-free and can be seamlessly applied to various MLLMs. Extensive experiments on multiple mainstream vision-language benchmarks demonstrate that HAWK achieves state-of-the-art accuracy. When applied to Qwen2.5-VL, HAWK retains 96.0% of the original accuracy after pruning 80.2% of the visual tokens. Additionally, it reduces end-to-end latency to 74.4% of the original and further decreases GPU memory usage across the tested models. The code is available at https://github.com/peppery77/HAWK.git.
title HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.07812