GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Mingzhe, Wang, Weijun, Ding, Xin, Mi, Liang, Wen, Hao, Li, Yuanchun, Pang, Lichen, Yang, Shansong, Liu, Yunxin, Cao, Ting
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911680312639488
author Huang, Mingzhe
Wang, Weijun
Ding, Xin
Mi, Liang
Wen, Hao
Li, Yuanchun
Pang, Lichen
Yang, Shansong
Liu, Yunxin
Cao, Ting
author_facet Huang, Mingzhe
Wang, Weijun
Ding, Xin
Mi, Liang
Wen, Hao
Li, Yuanchun
Pang, Lichen
Yang, Shansong
Liu, Yunxin
Cao, Ting
contents In Vision-Language Models (VLMs), processing a massive number of visual tokens incurs prohibitive computational overhead. While recent training-aware pruning methods attempt to selectively discard redundant tokens, they largely rely on continuous-gradient relaxations. However, visual token pruning is inherently a discrete, non-convex combinatorial problem; consequently, these continuous approximations frequently trap the optimization in sub-optimal local minima, especially under aggressive compression budgets. To overcome this fundamental bottleneck, we propose GRIP-VLM, a Group-Relative Importance Pruning framework driven by Reinforcement Learning. Rather than relying on smooth-gradient assumptions, GRIP-VLM formulates pruning as a Markov Decision Process, employing a Group Relative Policy Optimization (GRPO) paradigm anchored by supervised warm-up to directly explore the discrete selection space. Integrated with a budget-aware scorer, our lightweight agent dynamically evaluates per-token importance and adapts to arbitrary compression ratios without retraining. Extensive experiments across diverse multimodal benchmarks demonstrate that GRIP-VLM consistently outperforms heuristic and supervised-learning baselines, achieving a superior Pareto frontier and delivering up to a 15\% inference speedup at equal accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13375
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
Huang, Mingzhe
Wang, Weijun
Ding, Xin
Mi, Liang
Wen, Hao
Li, Yuanchun
Pang, Lichen
Yang, Shansong
Liu, Yunxin
Cao, Ting
Computer Vision and Pattern Recognition
Artificial Intelligence
In Vision-Language Models (VLMs), processing a massive number of visual tokens incurs prohibitive computational overhead. While recent training-aware pruning methods attempt to selectively discard redundant tokens, they largely rely on continuous-gradient relaxations. However, visual token pruning is inherently a discrete, non-convex combinatorial problem; consequently, these continuous approximations frequently trap the optimization in sub-optimal local minima, especially under aggressive compression budgets. To overcome this fundamental bottleneck, we propose GRIP-VLM, a Group-Relative Importance Pruning framework driven by Reinforcement Learning. Rather than relying on smooth-gradient assumptions, GRIP-VLM formulates pruning as a Markov Decision Process, employing a Group Relative Policy Optimization (GRPO) paradigm anchored by supervised warm-up to directly explore the discrete selection space. Integrated with a budget-aware scorer, our lightweight agent dynamically evaluates per-token importance and adapts to arbitrary compression ratios without retraining. Extensive experiments across diverse multimodal benchmarks demonstrate that GRIP-VLM consistently outperforms heuristic and supervised-learning baselines, achieving a superior Pareto frontier and delivering up to a 15\% inference speedup at equal accuracy.
title GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.13375