VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Sihan, Xu, Runsen, Cui, Chenhang, Wang, Tai, Lin, Dahua, Pang, Jiangmiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911149771980800
author Yang, Sihan
Xu, Runsen
Cui, Chenhang
Wang, Tai
Lin, Dahua
Pang, Jiangmiao
author_facet Yang, Sihan
Xu, Runsen
Cui, Chenhang
Wang, Tai
Lin, Dahua
Pang, Jiangmiao
contents Large Multimodal Models (LMMs) excel in visual-language tasks by leveraging numerous visual tokens for fine-grained visual information, but this token redundancy results in significant computational costs. Previous research aimed at reducing visual tokens during inference typically leverages importance maps derived from attention scores among vision-only tokens or vision-language tokens to prune tokens across one or multiple pruning stages. Despite this progress, pruning frameworks and strategies remain simplistic and insufficiently explored, often resulting in substantial performance degradation. In this paper, we propose VFlowOpt, a token pruning framework that introduces an importance map derivation process and a progressive pruning module with a recycling mechanism. The hyperparameters of its pruning strategy are further optimized by a visual information flow-guided method. Specifically, we compute an importance map for image tokens based on their attention-derived context relevance and patch-level information entropy. We then decide which tokens to retain or prune and aggregate the pruned ones as recycled tokens to avoid potential information loss. Finally, we apply a visual information flow-guided method that regards the last token in the LMM as the most representative signal of text-visual interactions. This method minimizes the discrepancy between token representations in LMMs with and without pruning, thereby enabling superior pruning strategies tailored to different LMMs. Experiments demonstrate that VFlowOpt can prune 90% of visual tokens while maintaining comparable performance, leading to an 89% reduction in KV-Cache memory and 3.8 times faster inference.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05211
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
Yang, Sihan
Xu, Runsen
Cui, Chenhang
Wang, Tai
Lin, Dahua
Pang, Jiangmiao
Computer Vision and Pattern Recognition
Large Multimodal Models (LMMs) excel in visual-language tasks by leveraging numerous visual tokens for fine-grained visual information, but this token redundancy results in significant computational costs. Previous research aimed at reducing visual tokens during inference typically leverages importance maps derived from attention scores among vision-only tokens or vision-language tokens to prune tokens across one or multiple pruning stages. Despite this progress, pruning frameworks and strategies remain simplistic and insufficiently explored, often resulting in substantial performance degradation. In this paper, we propose VFlowOpt, a token pruning framework that introduces an importance map derivation process and a progressive pruning module with a recycling mechanism. The hyperparameters of its pruning strategy are further optimized by a visual information flow-guided method. Specifically, we compute an importance map for image tokens based on their attention-derived context relevance and patch-level information entropy. We then decide which tokens to retain or prune and aggregate the pruned ones as recycled tokens to avoid potential information loss. Finally, we apply a visual information flow-guided method that regards the last token in the LMM as the most representative signal of text-visual interactions. This method minimizes the discrepancy between token representations in LMMs with and without pruning, thereby enabling superior pruning strategies tailored to different LMMs. Experiments demonstrate that VFlowOpt can prune 90% of visual tokens while maintaining comparable performance, leading to an 89% reduction in KV-Cache memory and 3.8 times faster inference.
title VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05211