How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, YiJie, Zhang, Yiqun, Jia, Zhuoyue, Yang, Xiaocui, Huang, Junzhao, Wang, Zihan, Feng, Shi, Wang, Daling, Zhang, Yifei, Liu, Yongkang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914571146493952
author Huang, YiJie
Zhang, Yiqun
Jia, Zhuoyue
Yang, Xiaocui
Huang, Junzhao
Wang, Zihan
Feng, Shi
Wang, Daling
Zhang, Yifei
Liu, Yongkang
author_facet Huang, YiJie
Zhang, Yiqun
Jia, Zhuoyue
Yang, Xiaocui
Huang, Junzhao
Wang, Zihan
Feng, Shi
Wang, Daling
Zhang, Yifei
Liu, Yongkang
contents Vision-language models improve perception by feeding increasingly long visual token sequences into language backbones, but the resulting inference cost raises a basic scaling question: as multimodal models grow, how many visual tokens are actually needed, and how should they be allocated under a fixed visual token budget? Existing training-free pruning methods typically answer this with one-shot proxies such as decoder attention, visual similarity, or conditional diversity. We argue that visual token pruning is better viewed as task-conditioned evidence search, especially under aggressive compression and across model scales. We propose F^3A, a training-free router for visual token pruning that operates before the language model consumes image tokens. F^3A builds lightweight question-conditioned cues, matches them to visual-grid tokens through frozen sparse sensing heads, and allocates a fixed vision token budget via coarse evidence localization, local refinement, coverage-preserving competition, and recovery of under-covered regions. It requires no model training, no extra LLM forward pass and preserves the original multimodal prompting and decoding pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16359
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
Huang, YiJie
Zhang, Yiqun
Jia, Zhuoyue
Yang, Xiaocui
Huang, Junzhao
Wang, Zihan
Feng, Shi
Wang, Daling
Zhang, Yifei
Liu, Yongkang
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-language models improve perception by feeding increasingly long visual token sequences into language backbones, but the resulting inference cost raises a basic scaling question: as multimodal models grow, how many visual tokens are actually needed, and how should they be allocated under a fixed visual token budget? Existing training-free pruning methods typically answer this with one-shot proxies such as decoder attention, visual similarity, or conditional diversity. We argue that visual token pruning is better viewed as task-conditioned evidence search, especially under aggressive compression and across model scales. We propose F^3A, a training-free router for visual token pruning that operates before the language model consumes image tokens. F^3A builds lightweight question-conditioned cues, matches them to visual-grid tokens through frozen sparse sensing heads, and allocates a fixed vision token budget via coarse evidence localization, local refinement, coverage-preserving competition, and recovery of under-covered regions. It requires no model training, no extra LLM forward pass and preserves the original multimodal prompting and decoding pipeline.
title How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.16359