TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hao, Lyu, Mengsi, He, Chenrui, Ao, Yulong, Lin, Yonghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918152845131776
author Zhang, Hao
Lyu, Mengsi
He, Chenrui
Ao, Yulong
Lin, Yonghua
author_facet Zhang, Hao
Lyu, Mengsi
He, Chenrui
Ao, Yulong
Lin, Yonghua
contents Large Multimodal Models (LMMs) have achieved significant success across various tasks. These models usually encode visual inputs into dense token sequences, which are then concatenated with textual tokens and jointly processed by a language model. However, the increased token count substantially raises computational and memory costs during inference. Token pruning has emerged as a promising approach to address this issue. Existing token pruning methods often rely on costly calibration or suboptimal importance metrics, leading to redundant retained tokens. In this paper, we analyze the redundancy differences between visual and textual tokens and propose pruning exclusively on visual tokens. Based on this, we propose a visual token pruning strategy that explicitly preserves both cross-modal alignment and intra-modal informational diversity. We introduce a mutual information-based token pruning strategy that removes visual tokens semantically misaligned with textual tokens, effectively preserving the alignment between the visual and textual modalities. To further improve the representational quality of the retained tokens, we additionally prune redundant visual tokens by maximizing the expected pairwise distances in the embedding space, which is solved efficiently with a greedy algorithm. Extensive experiments demonstrate that our method maintains strong performance while reducing tokens by 88.9% on models such as LLaVA-1.5-7B and LLaVA-NEXT-7B, resulting in a 56.7% improvement in inference speed.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00320
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
Zhang, Hao
Lyu, Mengsi
He, Chenrui
Ao, Yulong
Lin, Yonghua
Computer Vision and Pattern Recognition
Large Multimodal Models (LMMs) have achieved significant success across various tasks. These models usually encode visual inputs into dense token sequences, which are then concatenated with textual tokens and jointly processed by a language model. However, the increased token count substantially raises computational and memory costs during inference. Token pruning has emerged as a promising approach to address this issue. Existing token pruning methods often rely on costly calibration or suboptimal importance metrics, leading to redundant retained tokens. In this paper, we analyze the redundancy differences between visual and textual tokens and propose pruning exclusively on visual tokens. Based on this, we propose a visual token pruning strategy that explicitly preserves both cross-modal alignment and intra-modal informational diversity. We introduce a mutual information-based token pruning strategy that removes visual tokens semantically misaligned with textual tokens, effectively preserving the alignment between the visual and textual modalities. To further improve the representational quality of the retained tokens, we additionally prune redundant visual tokens by maximizing the expected pairwise distances in the embedding space, which is solved efficiently with a greedy algorithm. Extensive experiments demonstrate that our method maintains strong performance while reducing tokens by 88.9% on models such as LLaVA-1.5-7B and LLaVA-NEXT-7B, resulting in a 56.7% improvement in inference speed.
title TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00320