VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Hanxun, Li, Wentong, Qu, Xuan, Wang, Song, Chen, Junbo, Zhu, Jianke
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917365387624448
author Yu, Hanxun
Li, Wentong
Qu, Xuan
Wang, Song
Chen, Junbo
Zhu, Jianke
author_facet Yu, Hanxun
Li, Wentong
Qu, Xuan
Wang, Song
Chen, Junbo
Zhu, Jianke
contents Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high-resolution and video-based scenarios. Existing token reduction methods typically focus on isolated pipeline components and often neglect textual alignment, leading to performance degradation. In this paper, we propose VisionTrim, a unified framework for training-free MLLM acceleration, integrating two effective plug-and-play modules: 1) the Dominant Vision Token Selection (DVTS) module, which preserves essential visual tokens via a global-local view, and 2) the Text-Guided Vision Complement (TGVC) module, which facilitates context-aware token merging guided by textual cues. Extensive experiments across diverse image and video multimodal benchmarks demonstrate the performance superiority of our VisionTrim, advancing practical MLLM deployment in real-world applications. The code is available at: https://github.com/hanxunyu/VisionTrim.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22674
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
Yu, Hanxun
Li, Wentong
Qu, Xuan
Wang, Song
Chen, Junbo
Zhu, Jianke
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high-resolution and video-based scenarios. Existing token reduction methods typically focus on isolated pipeline components and often neglect textual alignment, leading to performance degradation. In this paper, we propose VisionTrim, a unified framework for training-free MLLM acceleration, integrating two effective plug-and-play modules: 1) the Dominant Vision Token Selection (DVTS) module, which preserves essential visual tokens via a global-local view, and 2) the Text-Guided Vision Complement (TGVC) module, which facilitates context-aware token merging guided by textual cues. Extensive experiments across diverse image and video multimodal benchmarks demonstrate the performance superiority of our VisionTrim, advancing practical MLLM deployment in real-world applications. The code is available at: https://github.com/hanxunyu/VisionTrim.
title VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.22674