Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhan, Wengyi, Lin, Mingbao, Lin, Zhihang, Ji, Rongrong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914168689393664
author Zhan, Wengyi
Lin, Mingbao
Lin, Zhihang
Ji, Rongrong
author_facet Zhan, Wengyi
Lin, Mingbao
Lin, Zhihang
Ji, Rongrong
contents Multimodal large language models (MLLMs) deliver impressive vision-language reasoning but suffer steep inference latency because self-attention scales quadratically with sequence length and thousands of visual tokens contributed by high-resolution images. Naively pruning less-informative visual tokens reduces this burden, yet indiscriminate removal can strip away contextual cues essential for background or fine-grained questions, undermining accuracy. In this paper, we present ParVTS (Parallel Vision Token Scheduling), a training-free scheduling framework that partitions visual tokens into subject and non-subject groups, processes them in parallel to transfer their semantics into question tokens, and discards the non-subject path mid-inference to reduce computation. This scheduling reduces computational complexity, requires no heuristics or additional modules, and is compatible with diverse existing MLLM architectures. Experiments across multiple MLLM backbones show that ParVTS prunes up to 88.9% of visual tokens with minimal performance drop, achieving 1.77x speedup and 70% FLOPs reduction.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
Zhan, Wengyi
Lin, Mingbao
Lin, Zhihang
Ji, Rongrong
Computer Vision and Pattern Recognition
Multimedia
Multimodal large language models (MLLMs) deliver impressive vision-language reasoning but suffer steep inference latency because self-attention scales quadratically with sequence length and thousands of visual tokens contributed by high-resolution images. Naively pruning less-informative visual tokens reduces this burden, yet indiscriminate removal can strip away contextual cues essential for background or fine-grained questions, undermining accuracy. In this paper, we present ParVTS (Parallel Vision Token Scheduling), a training-free scheduling framework that partitions visual tokens into subject and non-subject groups, processes them in parallel to transfer their semantics into question tokens, and discards the non-subject path mid-inference to reduce computation. This scheduling reduces computational complexity, requires no heuristics or additional modules, and is compatible with diverse existing MLLM architectures. Experiments across multiple MLLM backbones show that ParVTS prunes up to 88.9% of visual tokens with minimal performance drop, achieving 1.77x speedup and 70% FLOPs reduction.
title Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2511.18875