OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Morunliu, Xu, Ruotao, Li, Le, Wang, Yue, Zhang, Jianxin, Li, Juntao, Lou, Yihang, Feng, Siwei, Li, Peifeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913140500856832
author Yang, Morunliu
Xu, Ruotao
Li, Le
Wang, Yue
Zhang, Jianxin
Li, Juntao
Lou, Yihang
Feng, Siwei
Li, Peifeng
author_facet Yang, Morunliu
Xu, Ruotao
Li, Le
Wang, Yue
Zhang, Jianxin
Li, Juntao
Lou, Yihang
Feng, Siwei
Li, Peifeng
contents Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences introduces substantial computational overhead, making efficient token compression crucial. Existing methods typically rely on fixed, modality-specific guidance, which fails to account for the varying importance of modalities across different queries. To address this limitation, we propose $\textbf{OmniSelect}$, a training-free, modality-adaptive token pruning framework that dynamically selects appropriate compression strategies for multimodal inputs. Specifically, we leverage a lightweight AudioCLIP model to estimate cross-modal relevance and categorize each input into three pruning regimes: Audio-Centric, Video-Centric, and Uniform pruning. Based on these relevance scores, OmniSelect further performs fine-grained token pruning within each temporal group, adaptively allocating pruning ratios to preserve informative tokens across modalities. By explicitly modeling modality preference and enabling dynamic strategy selection, OmniSelect effectively avoids the pitfalls of one-size-fits-all compression. Extensive experiments demonstrate that our method achieves efficient multimodal token reduction while maintaining strong performance, without requiring any additional training.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18041
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
Yang, Morunliu
Xu, Ruotao
Li, Le
Wang, Yue
Zhang, Jianxin
Li, Juntao
Lou, Yihang
Feng, Siwei
Li, Peifeng
Computer Vision and Pattern Recognition
Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences introduces substantial computational overhead, making efficient token compression crucial. Existing methods typically rely on fixed, modality-specific guidance, which fails to account for the varying importance of modalities across different queries. To address this limitation, we propose $\textbf{OmniSelect}$, a training-free, modality-adaptive token pruning framework that dynamically selects appropriate compression strategies for multimodal inputs. Specifically, we leverage a lightweight AudioCLIP model to estimate cross-modal relevance and categorize each input into three pruning regimes: Audio-Centric, Video-Centric, and Uniform pruning. Based on these relevance scores, OmniSelect further performs fine-grained token pruning within each temporal group, adaptively allocating pruning ratios to preserve informative tokens across modalities. By explicitly modeling modality preference and enabling dynamic strategy selection, OmniSelect effectively avoids the pitfalls of one-size-fits-all compression. Extensive experiments demonstrate that our method achieves efficient multimodal token reduction while maintaining strong performance, without requiring any additional training.
title OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18041