ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Pu, Li, Yuwei, Xian, Xingyuan, Tang, Guoming
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909857446100992
author Zhang, Pu
Li, Yuwei
Xian, Xingyuan
Tang, Guoming
author_facet Zhang, Pu
Li, Yuwei
Xian, Xingyuan
Tang, Guoming
contents As the capabilities of Vision-Language Models (VLMs) advance, they can process increasingly large inputs, which, unlike in LLMs, generates significant visual token redundancy and leads to prohibitive inference costs. While many methods aim to reduce these costs by pruning visual tokens, existing approaches, whether based on attention or diversity, typically neglect the guidance of the text prompt and thus fail to prioritize task relevance. In this work, we propose a novel, zero-shot method that reframes the problem by introducing a prompt-aware perspective, explicitly modeling visual token pruning as a balance between task relevance and information diversity. Our hierarchical approach first selects a core set of task-relevant visual tokens and then supplements them with diversity tokens to preserve broader context. Experiments across multiple models and benchmarks show that our method achieves performance that matches or surpasses the state-of-the-art with only minimal accuracy loss, even when pruning up to 90\% of the tokens. Furthermore, these gains are accompanied by significant reductions in GPU memory footprint and inference latency.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17197
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
Zhang, Pu
Li, Yuwei
Xian, Xingyuan
Tang, Guoming
Computer Vision and Pattern Recognition
Artificial Intelligence
As the capabilities of Vision-Language Models (VLMs) advance, they can process increasingly large inputs, which, unlike in LLMs, generates significant visual token redundancy and leads to prohibitive inference costs. While many methods aim to reduce these costs by pruning visual tokens, existing approaches, whether based on attention or diversity, typically neglect the guidance of the text prompt and thus fail to prioritize task relevance. In this work, we propose a novel, zero-shot method that reframes the problem by introducing a prompt-aware perspective, explicitly modeling visual token pruning as a balance between task relevance and information diversity. Our hierarchical approach first selects a core set of task-relevant visual tokens and then supplements them with diversity tokens to preserve broader context. Experiments across multiple models and benchmarks show that our method achieves performance that matches or surpasses the state-of-the-art with only minimal accuracy loss, even when pruning up to 90\% of the tokens. Furthermore, these gains are accompanied by significant reductions in GPU memory footprint and inference latency.
title ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.17197