Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xuyang, Wang, Yiyu, Ma, Junpeng, Zhang, Linfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908661057585152
author Liu, Xuyang
Wang, Yiyu
Ma, Junpeng
Zhang, Linfeng
author_facet Liu, Xuyang
Wang, Yiyu
Ma, Junpeng
Zhang, Linfeng
contents Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token compression methods for VideoLLMs reveals two critical issues: (i) overlooking distinctive visual signals across frames, leading to information loss; (ii) suffering from implementation constraints, causing incompatibility with modern architectures or efficient operators. To address these challenges, we distill three design principles for VideoLLM token compression and propose a plug-and-play inference acceleration framework "Video Compression Commander" (VidCom2). By quantifying each frame's uniqueness, VidCom2 adaptively adjusts compression intensity across frames, effectively preserving essential information while reducing redundancy in video sequences. Extensive experiments across various VideoLLMs and benchmarks demonstrate the superior performance and efficiency of our VidCom2. With only 25% visual tokens, VidCom2 achieves 99.6% of the original performance on LLaVA-OV while reducing 70.8% of the LLM generation latency. Notably, our Frame Compression Adjustment strategy is compatible with other token compression methods to further improve their performance. Our code is available at https://github.com/xuyang-liu16/VidCom2.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
Liu, Xuyang
Wang, Yiyu
Ma, Junpeng
Zhang, Linfeng
Computer Vision and Pattern Recognition
Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token compression methods for VideoLLMs reveals two critical issues: (i) overlooking distinctive visual signals across frames, leading to information loss; (ii) suffering from implementation constraints, causing incompatibility with modern architectures or efficient operators. To address these challenges, we distill three design principles for VideoLLM token compression and propose a plug-and-play inference acceleration framework "Video Compression Commander" (VidCom2). By quantifying each frame's uniqueness, VidCom2 adaptively adjusts compression intensity across frames, effectively preserving essential information while reducing redundancy in video sequences. Extensive experiments across various VideoLLMs and benchmarks demonstrate the superior performance and efficiency of our VidCom2. With only 25% visual tokens, VidCom2 achieves 99.6% of the original performance on LLaVA-OV while reducing 70.8% of the LLM generation latency. Notably, our Frame Compression Adjustment strategy is compatible with other token compression methods to further improve their performance. Our code is available at https://github.com/xuyang-liu16/VidCom2.
title Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.14454