EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Jiafei, Zhou, Fengwei, Qu, Jin, Li, Wenjin Jason, Wu, Tong, Xue, Gengjian, Zhao, Zhikang, Wei, Daomin, Lu, Yichao, Na, Bailin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915943696826368
author Song, Jiafei
Zhou, Fengwei
Qu, Jin
Li, Wenjin Jason
Wu, Tong
Xue, Gengjian
Zhao, Zhikang
Wei, Daomin
Lu, Yichao
Na, Bailin
author_facet Song, Jiafei
Zhou, Fengwei
Qu, Jin
Li, Wenjin Jason
Wu, Tong
Xue, Gengjian
Zhao, Zhikang
Wei, Daomin
Lu, Yichao
Na, Bailin
contents Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficiency is often hampered by the large number of visual tokens, particularly in high-resolution or multi-image scenarios. To address this issue, we propose EvoComp, a visual token compression framework that significantly reduces token count while preserving task accuracy. EvoComp introduces a lightweight encoder-only transformer-based compressor that selects the most informative and non-redundant visual tokens by jointly considering visual and textual contexts. A core challenge lies in providing effective supervision for training the compressor. To this end, we design an evolutionary labeling strategy that searches for token subsets minimizing the MLLM's output loss, while enforcing semantic diversity through vocabulary-based token grouping. We further train the compressor using a tailored loss function combining the GHM loss to mitigate class and difficulty imbalance, and a cosine similarity regularization to encourage semantic separation between retained and discarded tokens. Extensive experiments across multiple vision-language benchmarks show that EvoComp outperforms existing methods based on attention or similarity heuristics. Notably, it retains 99.3% of the original accuracy under 3x token compression and delivers up to 1.6x speedup on mobile devices.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17087
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
Song, Jiafei
Zhou, Fengwei
Qu, Jin
Li, Wenjin Jason
Wu, Tong
Xue, Gengjian
Zhao, Zhikang
Wei, Daomin
Lu, Yichao
Na, Bailin
Computer Vision and Pattern Recognition
Machine Learning
Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficiency is often hampered by the large number of visual tokens, particularly in high-resolution or multi-image scenarios. To address this issue, we propose EvoComp, a visual token compression framework that significantly reduces token count while preserving task accuracy. EvoComp introduces a lightweight encoder-only transformer-based compressor that selects the most informative and non-redundant visual tokens by jointly considering visual and textual contexts. A core challenge lies in providing effective supervision for training the compressor. To this end, we design an evolutionary labeling strategy that searches for token subsets minimizing the MLLM's output loss, while enforcing semantic diversity through vocabulary-based token grouping. We further train the compressor using a tailored loss function combining the GHM loss to mitigate class and difficulty imbalance, and a cosine similarity regularization to encourage semantic separation between retained and discarded tokens. Extensive experiments across multiple vision-language benchmarks show that EvoComp outperforms existing methods based on attention or similarity heuristics. Notably, it retains 99.3% of the original accuracy under 3x token compression and delivers up to 1.6x speedup on mobile devices.
title EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2604.17087