Saved in:
Bibliographic Details
Main Authors: Yang, Zhongyu, Xu, Dannong, Pang, Wei, Yuan, Yingfang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.01949
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912741016469504
author Yang, Zhongyu
Xu, Dannong
Pang, Wei
Yuan, Yingfang
author_facet Yang, Zhongyu
Xu, Dannong
Pang, Wei
Yuan, Yingfang
contents The rapid growth of visual tokens in multimodal large language models (MLLMs) leads to excessive memory consumption and inference latency, especially when handling high-resolution images and videos. Token pruning is a technique used to mitigate this issue by removing redundancy, but existing methods often ignore relevance to the user query or suffer from the limitations of attention mechanisms, reducing their adaptability and effectiveness. To address these challenges, we propose Script, a plug-and-play pruning method that requires no retraining and generalizes across diverse MLLMs. Script comprises two modules: a graph-structured pruning module that removes visually redundant tokens, and a query-conditioned semantic pruning module that preserves query-relevant visual information. Together, they enhance performance on multimodal tasks. Experiments on fourteen benchmarks across image and video understanding tasks show that Script consistently achieves higher model efficiency and predictive accuracy compared to existing pruning methods. On LLaVA-NeXT-7B, it achieves up to 6.8x prefill speedup and 10x FLOP reduction, while retaining 96.88% of the original performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
Yang, Zhongyu
Xu, Dannong
Pang, Wei
Yuan, Yingfang
Computer Vision and Pattern Recognition
The rapid growth of visual tokens in multimodal large language models (MLLMs) leads to excessive memory consumption and inference latency, especially when handling high-resolution images and videos. Token pruning is a technique used to mitigate this issue by removing redundancy, but existing methods often ignore relevance to the user query or suffer from the limitations of attention mechanisms, reducing their adaptability and effectiveness. To address these challenges, we propose Script, a plug-and-play pruning method that requires no retraining and generalizes across diverse MLLMs. Script comprises two modules: a graph-structured pruning module that removes visually redundant tokens, and a query-conditioned semantic pruning module that preserves query-relevant visual information. Together, they enhance performance on multimodal tasks. Experiments on fourteen benchmarks across image and video understanding tasks show that Script consistently achieves higher model efficiency and predictive accuracy compared to existing pruning methods. On LLaVA-NeXT-7B, it achieves up to 6.8x prefill speedup and 10x FLOP reduction, while retaining 96.88% of the original performance.
title Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01949