OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Xiwen, Zhu, Wenhui, Li, Gen, Dong, Xuanzhao, Xiong, Yujian, Wang, Hao, Qiu, Peijie, Song, Qingquan, Wang, Zhipeng, Tang, Shao, Wang, Yalin, Razi, Abolfazl
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917376655622144
author Chen, Xiwen
Zhu, Wenhui
Li, Gen
Dong, Xuanzhao
Xiong, Yujian
Wang, Hao
Qiu, Peijie
Song, Qingquan
Wang, Zhipeng
Tang, Shao
Wang, Yalin
Razi, Abolfazl
author_facet Chen, Xiwen
Zhu, Wenhui
Li, Gen
Dong, Xuanzhao
Xiong, Yujian
Wang, Hao
Qiu, Peijie
Song, Qingquan
Wang, Zhipeng
Tang, Shao
Wang, Yalin
Razi, Abolfazl
contents Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual token pruning to accelerate inference, while existing pruning methods overlook the underlying distributional structure of visual representations. We propose OTPrune, a training-free framework that formulates pruning as distribution alignment via optimal transport (OT). By minimizing the 2-Wasserstein distance between the full and pruned token distributions, OTPrune preserves both local diversity and global representativeness while reducing inference cost. Moreover, we derive a tractable submodular objective that enables efficient optimization, and theoretically prove its monotonicity and submodularity, providing a principled foundation for stable and efficient pruning. We further provide a comprehensive analysis that explains how distributional alignment contributes to stable and semantically faithful pruning. Comprehensive experiments on wider benchmarks demonstrate that OTPrune achieves superior performance-efficiency tradeoffs compared to state-of-the-art methods. The code is available at https://github.com/xiwenc1/OTPrune.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20205
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
Chen, Xiwen
Zhu, Wenhui
Li, Gen
Dong, Xuanzhao
Xiong, Yujian
Wang, Hao
Qiu, Peijie
Song, Qingquan
Wang, Zhipeng
Tang, Shao
Wang, Yalin
Razi, Abolfazl
Computer Vision and Pattern Recognition
Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual token pruning to accelerate inference, while existing pruning methods overlook the underlying distributional structure of visual representations. We propose OTPrune, a training-free framework that formulates pruning as distribution alignment via optimal transport (OT). By minimizing the 2-Wasserstein distance between the full and pruned token distributions, OTPrune preserves both local diversity and global representativeness while reducing inference cost. Moreover, we derive a tractable submodular objective that enables efficient optimization, and theoretically prove its monotonicity and submodularity, providing a principled foundation for stable and efficient pruning. We further provide a comprehensive analysis that explains how distributional alignment contributes to stable and semantically faithful pruning. Comprehensive experiments on wider benchmarks demonstrate that OTPrune achieves superior performance-efficiency tradeoffs compared to state-of-the-art methods. The code is available at https://github.com/xiwenc1/OTPrune.
title OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.20205