Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Sihan, Zhang, Jianwei, Zheng, Pengcheng, Yan, Jiaxin, Qin, Caiyan, Ye, Yalan, Dong, Wei, Wang, Peng, Yang, Yang, Zhang, Chaoning
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917340655910912
author Cao, Sihan
Zhang, Jianwei
Zheng, Pengcheng
Yan, Jiaxin
Qin, Caiyan
Ye, Yalan
Dong, Wei
Wang, Peng
Yang, Yang
Zhang, Chaoning
author_facet Cao, Sihan
Zhang, Jianwei
Zheng, Pengcheng
Yan, Jiaxin
Qin, Caiyan
Ye, Yalan
Dong, Wei
Wang, Peng
Yang, Yang
Zhang, Chaoning
contents Large Vision-Language Models (LVLMs) incur substantial inference costs due to the processing of a vast number of visual tokens. Existing methods typically struggle to model progressive visual token reduction as a multi-step decision process with sequential dependencies and often rely on hand-engineered scoring rules that lack adaptive optimization for complex reasoning trajectories. To overcome these limitations, we propose TPRL, a reinforcement learning framework that learns adaptive pruning trajectories through language-guided sequential optimization tied directly to end-task performance. We formulate visual token pruning as a sequential decision process with explicit state transitions and employ a self-supervised autoencoder to compress visual tokens into a compact state representation for efficient policy learning. The pruning policy is initialized through learning from demonstrations and subsequently fine-tuned using Proximal Policy Optimization (PPO) to jointly optimize task accuracy and computational efficiency. Our experimental results demonstrate that TPRL removes up to 66.7\% of visual tokens and achieves up to a 54.2\% reduction in FLOPs during inference while maintaining a near-lossless average accuracy drop of only 0.7\%. Code is released at \href{https://github.com/MagicVicCoder/TPRL}{\textcolor{mypink}{https://github.com/MagicVicCoder/TPRL}}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13394
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
Cao, Sihan
Zhang, Jianwei
Zheng, Pengcheng
Yan, Jiaxin
Qin, Caiyan
Ye, Yalan
Dong, Wei
Wang, Peng
Yang, Yang
Zhang, Chaoning
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) incur substantial inference costs due to the processing of a vast number of visual tokens. Existing methods typically struggle to model progressive visual token reduction as a multi-step decision process with sequential dependencies and often rely on hand-engineered scoring rules that lack adaptive optimization for complex reasoning trajectories. To overcome these limitations, we propose TPRL, a reinforcement learning framework that learns adaptive pruning trajectories through language-guided sequential optimization tied directly to end-task performance. We formulate visual token pruning as a sequential decision process with explicit state transitions and employ a self-supervised autoencoder to compress visual tokens into a compact state representation for efficient policy learning. The pruning policy is initialized through learning from demonstrations and subsequently fine-tuned using Proximal Policy Optimization (PPO) to jointly optimize task accuracy and computational efficiency. Our experimental results demonstrate that TPRL removes up to 66.7\% of visual tokens and achieves up to a 54.2\% reduction in FLOPs during inference while maintaining a near-lossless average accuracy drop of only 0.7\%. Code is released at \href{https://github.com/MagicVicCoder/TPRL}{\textcolor{mypink}{https://github.com/MagicVicCoder/TPRL}}.
title Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13394