Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Landi, Yao, Mingde, Young, Shawn, Xu, Lijian
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911723661819904
author He, Landi
Yao, Mingde
Young, Shawn
Xu, Lijian
author_facet He, Landi
Yao, Mingde
Young, Shawn
Xu, Lijian
contents Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during training. However, the optimization is driven by surrogate gradients rather than the true selection process, leading to unreliable learning of token importance. In this paper, we propose DiffPrune, which reformulates pruning as continuous control of token information instead of discrete selection learning. Specifically, we introduce an Information Throttler that modulates each token using variance-preserving noise conditioned on importance scores, where higher scores induce less information suppression during training. This design directly operates on token representations, naturally providing a fully differentiable optimization path for learning token importance. At inference, tokens are removed via hard thresholding on the learned scores. Across ten VLM benchmarks, DiffPrune retains 96.5% of full-model accuracy while accelerating LLM prefill by 2.85x, with only 0.69 ms of inference overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28051
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
He, Landi
Yao, Mingde
Young, Shawn
Xu, Lijian
Computer Vision and Pattern Recognition
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during training. However, the optimization is driven by surrogate gradients rather than the true selection process, leading to unreliable learning of token importance. In this paper, we propose DiffPrune, which reformulates pruning as continuous control of token information instead of discrete selection learning. Specifically, we introduce an Information Throttler that modulates each token using variance-preserving noise conditioned on importance scores, where higher scores induce less information suppression during training. This design directly operates on token representations, naturally providing a fully differentiable optimization path for learning token importance. At inference, tokens are removed via hard thresholding on the learned scores. Across ten VLM benchmarks, DiffPrune retains 96.5% of full-model accuracy while accelerating LLM prefill by 2.85x, with only 0.69 ms of inference overhead.
title Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.28051