Task-Aware KV Compression For Cost-Effective Long Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Minghao, Shu, Yan, Zhang, Peitian, Lun, Kun, Yuan, Huaying, Zhou, Juenjie, Xiao, Shitao, Zhao, Bo, Liu, Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918071147429888
author Qin, Minghao
Shu, Yan
Zhang, Peitian
Lun, Kun
Yuan, Huaying
Zhou, Juenjie
Xiao, Shitao
Zhao, Bo
Liu, Zheng
author_facet Qin, Minghao
Shu, Yan
Zhang, Peitian
Lun, Kun
Yuan, Huaying
Zhou, Juenjie
Xiao, Shitao
Zhao, Bo
Liu, Zheng
contents Long-video understanding (LVU) remains a severe challenge for existing multimodal large language models (MLLMs), primarily due to the prohibitive computational cost. Recent approaches have explored KV compression to mitigate this issue, but they often suffer from significant information loss at high compression ratios. In this paper, we introduce Video-X^2L, which flexibly preserves critical video information for each LVU task. Video-X^2L involves two key operations. The first one is called bi-level KV compression. During the MLLM's pre-filling stage, Video-X^2L generates two types of compressed KVs: low-compression KVs (L-KVs) to capture fine-grained video details and high-compression KVs (H-KVs) to offer compact video representations. The second one is called selective KV re-loading. During the MLLM's decoding stage, Video-X^2L selectively re-loads L-KVs for the most critical video chunks while using H-KVs for other less important ones. This allows the MLLM to fully utilize task-specific information while maintaining the overall compactness. Video-X^2L is simple yet effective: it is free from additional training and directly compatible with existing KV-compressible MLLMs. We evaluate Video-X^2L with a variety of popular LVU benchmarks, including VideoMME, MLVU, LongVideoBench, and VNBench. Our experiment result shows that Video-X^2L outperforms existing KV-compression methods by a huge advantage while substantially saving the computation cost.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Task-Aware KV Compression For Cost-Effective Long Video Understanding
Qin, Minghao
Shu, Yan
Zhang, Peitian
Lun, Kun
Yuan, Huaying
Zhou, Juenjie
Xiao, Shitao
Zhao, Bo
Liu, Zheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Long-video understanding (LVU) remains a severe challenge for existing multimodal large language models (MLLMs), primarily due to the prohibitive computational cost. Recent approaches have explored KV compression to mitigate this issue, but they often suffer from significant information loss at high compression ratios. In this paper, we introduce Video-X^2L, which flexibly preserves critical video information for each LVU task. Video-X^2L involves two key operations. The first one is called bi-level KV compression. During the MLLM's pre-filling stage, Video-X^2L generates two types of compressed KVs: low-compression KVs (L-KVs) to capture fine-grained video details and high-compression KVs (H-KVs) to offer compact video representations. The second one is called selective KV re-loading. During the MLLM's decoding stage, Video-X^2L selectively re-loads L-KVs for the most critical video chunks while using H-KVs for other less important ones. This allows the MLLM to fully utilize task-specific information while maintaining the overall compactness. Video-X^2L is simple yet effective: it is free from additional training and directly compatible with existing KV-compressible MLLMs. We evaluate Video-X^2L with a variety of popular LVU benchmarks, including VideoMME, MLVU, LongVideoBench, and VNBench. Our experiment result shows that Video-X^2L outperforms existing KV-compression methods by a huge advantage while substantially saving the computation cost.
title Task-Aware KV Compression For Cost-Effective Long Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.21184