VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917365387624448 |
|---|---|
| author | Yu, Hanxun Li, Wentong Qu, Xuan Wang, Song Chen, Junbo Zhu, Jianke |
| author_facet | Yu, Hanxun Li, Wentong Qu, Xuan Wang, Song Chen, Junbo Zhu, Jianke |
| contents | Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high-resolution and video-based scenarios. Existing token reduction methods typically focus on isolated pipeline components and often neglect textual alignment, leading to performance degradation. In this paper, we propose VisionTrim, a unified framework for training-free MLLM acceleration, integrating two effective plug-and-play modules: 1) the Dominant Vision Token Selection (DVTS) module, which preserves essential visual tokens via a global-local view, and 2) the Text-Guided Vision Complement (TGVC) module, which facilitates context-aware token merging guided by textual cues. Extensive experiments across diverse image and video multimodal benchmarks demonstrate the performance superiority of our VisionTrim, advancing practical MLLM deployment in real-world applications. The code is available at: https://github.com/hanxunyu/VisionTrim. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_22674 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration Yu, Hanxun Li, Wentong Qu, Xuan Wang, Song Chen, Junbo Zhu, Jianke Computer Vision and Pattern Recognition Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high-resolution and video-based scenarios. Existing token reduction methods typically focus on isolated pipeline components and often neglect textual alignment, leading to performance degradation. In this paper, we propose VisionTrim, a unified framework for training-free MLLM acceleration, integrating two effective plug-and-play modules: 1) the Dominant Vision Token Selection (DVTS) module, which preserves essential visual tokens via a global-local view, and 2) the Text-Guided Vision Complement (TGVC) module, which facilitates context-aware token merging guided by textual cues. Extensive experiments across diverse image and video multimodal benchmarks demonstrate the performance superiority of our VisionTrim, advancing practical MLLM deployment in real-world applications. The code is available at: https://github.com/hanxunyu/VisionTrim. |
| title | VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.22674 |