TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Xudong, Ye, Peng, Tu, Chongjun, Cao, Jianjian, Yang, Yaoxin, Zhang, Lin, Zhou, Dongzhan, Chen, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910874140147712
author Tan, Xudong
Ye, Peng
Tu, Chongjun
Cao, Jianjian
Yang, Yaoxin
Zhang, Lin
Zhou, Dongzhan
Chen, Tao
author_facet Tan, Xudong
Ye, Peng
Tu, Chongjun
Cao, Jianjian
Yang, Yaoxin
Zhang, Lin
Zhou, Dongzhan
Chen, Tao
contents Multimodal Large Language Models (MLLMs) are becoming increasingly popular, while the high computational cost associated with multimodal data input, particularly from visual tokens, poses a significant challenge. Existing training-based token compression methods improve inference efficiency but require costly retraining, while training-free methods struggle to maintain performance when aggressively reducing token counts. In this study, we reveal that the performance degradation of MLLM closely correlates with the accelerated loss of information in the attention output matrix. This insight introduces a novel information-preserving perspective, making it possible to maintain performance even under extreme token compression. Based on this finding, we propose TokenCarve, a training-free, plug-and-play, two-stage token compression framework. The first stage employs an Information-Preservation-Guided Selection (IPGS) strategy to prune low-information tokens, while the second stage further leverages IPGS to guide token merging, minimizing information loss. Extensive experiments on 11 datasets and 2 model variants demonstrate the effectiveness of TokenCarve. It can even reduce the number of visual tokens to 22.2% of the original count, achieving a 1.23x speedup in inference, a 64% reduction in KV cache storage, and only a 1.54% drop in accuracy. Our code is available at https://github.com/ShawnTan86/TokenCarve.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
Tan, Xudong
Ye, Peng
Tu, Chongjun
Cao, Jianjian
Yang, Yaoxin
Zhang, Lin
Zhou, Dongzhan
Chen, Tao
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) are becoming increasingly popular, while the high computational cost associated with multimodal data input, particularly from visual tokens, poses a significant challenge. Existing training-based token compression methods improve inference efficiency but require costly retraining, while training-free methods struggle to maintain performance when aggressively reducing token counts. In this study, we reveal that the performance degradation of MLLM closely correlates with the accelerated loss of information in the attention output matrix. This insight introduces a novel information-preserving perspective, making it possible to maintain performance even under extreme token compression. Based on this finding, we propose TokenCarve, a training-free, plug-and-play, two-stage token compression framework. The first stage employs an Information-Preservation-Guided Selection (IPGS) strategy to prune low-information tokens, while the second stage further leverages IPGS to guide token merging, minimizing information loss. Extensive experiments on 11 datasets and 2 model variants demonstrate the effectiveness of TokenCarve. It can even reduce the number of visual tokens to 22.2% of the original count, achieving a 1.23x speedup in inference, a 64% reduction in KV cache storage, and only a 1.54% drop in accuracy. Our code is available at https://github.com/ShawnTan86/TokenCarve.
title TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10501