How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914571146493952 |
|---|---|
| author | Huang, YiJie Zhang, Yiqun Jia, Zhuoyue Yang, Xiaocui Huang, Junzhao Wang, Zihan Feng, Shi Wang, Daling Zhang, Yifei Liu, Yongkang |
| author_facet | Huang, YiJie Zhang, Yiqun Jia, Zhuoyue Yang, Xiaocui Huang, Junzhao Wang, Zihan Feng, Shi Wang, Daling Zhang, Yifei Liu, Yongkang |
| contents | Vision-language models improve perception by feeding increasingly long visual token sequences into language backbones, but the resulting inference cost raises a basic scaling question: as multimodal models grow, how many visual tokens are actually needed, and how should they be allocated under a fixed visual token budget? Existing training-free pruning methods typically answer this with one-shot proxies such as decoder attention, visual similarity, or conditional diversity. We argue that visual token pruning is better viewed as task-conditioned evidence search, especially under aggressive compression and across model scales. We propose F^3A, a training-free router for visual token pruning that operates before the language model consumes image tokens. F^3A builds lightweight question-conditioned cues, matches them to visual-grid tokens through frozen sparse sensing heads, and allocates a fixed vision token budget via coarse evidence localization, local refinement, coverage-preserving competition, and recovery of under-covered regions. It requires no model training, no extra LLM forward pass and preserves the original multimodal prompting and decoding pipeline. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_16359 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A Huang, YiJie Zhang, Yiqun Jia, Zhuoyue Yang, Xiaocui Huang, Junzhao Wang, Zihan Feng, Shi Wang, Daling Zhang, Yifei Liu, Yongkang Computer Vision and Pattern Recognition Artificial Intelligence Vision-language models improve perception by feeding increasingly long visual token sequences into language backbones, but the resulting inference cost raises a basic scaling question: as multimodal models grow, how many visual tokens are actually needed, and how should they be allocated under a fixed visual token budget? Existing training-free pruning methods typically answer this with one-shot proxies such as decoder attention, visual similarity, or conditional diversity. We argue that visual token pruning is better viewed as task-conditioned evidence search, especially under aggressive compression and across model scales. We propose F^3A, a training-free router for visual token pruning that operates before the language model consumes image tokens. F^3A builds lightweight question-conditioned cues, matches them to visual-grid tokens through frozen sparse sensing heads, and allocates a fixed vision token budget via coarse evidence localization, local refinement, coverage-preserving competition, and recovery of under-covered regions. It requires no model training, no extra LLM forward pass and preserves the original multimodal prompting and decoding pipeline. |
| title | How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2605.16359 |