AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiao, Si, Qingyi, Wu, Jianlong, Zhu, Shiyu, Cao, Li, Nie, Liqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909642184982528
author Wang, Xiao
Si, Qingyi
Wu, Jianlong
Zhu, Shiyu
Cao, Li
Nie, Liqiang
author_facet Wang, Xiao
Si, Qingyi
Wu, Jianlong
Zhu, Shiyu
Cao, Li
Nie, Liqiang
contents Multimodal Large Language Models (MLLMs) have revolutionized video understanding, yet are still limited by context length when processing long videos. Recent methods compress videos by leveraging visual redundancy uniformly, yielding promising results. Nevertheless, our quantitative analysis shows that redundancy varies significantly across time and model layers, necessitating a more flexible compression strategy. We propose AdaReTaKe, a training-free method that flexibly reduces visual redundancy by allocating compression ratios among time and layers with theoretical guarantees. Integrated into state-of-the-art MLLMs, AdaReTaKe improves processing capacity from 256 to 2048 frames while preserving critical information. Experiments on VideoMME, MLVU, LongVideoBench, and LVBench datasets demonstrate that AdaReTaKe outperforms existing methods by 2.3% and 2.8% for 7B and 72B models, respectively, with even greater improvements of 5.9% and 6.0% on the longest LVBench. Our code is available at https://github.com/SCZwangxiao/video-FlexReduc.git.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
Wang, Xiao
Si, Qingyi
Wu, Jianlong
Zhu, Shiyu
Cao, Li
Nie, Liqiang
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Multimodal Large Language Models (MLLMs) have revolutionized video understanding, yet are still limited by context length when processing long videos. Recent methods compress videos by leveraging visual redundancy uniformly, yielding promising results. Nevertheless, our quantitative analysis shows that redundancy varies significantly across time and model layers, necessitating a more flexible compression strategy. We propose AdaReTaKe, a training-free method that flexibly reduces visual redundancy by allocating compression ratios among time and layers with theoretical guarantees. Integrated into state-of-the-art MLLMs, AdaReTaKe improves processing capacity from 256 to 2048 frames while preserving critical information. Experiments on VideoMME, MLVU, LongVideoBench, and LVBench datasets demonstrate that AdaReTaKe outperforms existing methods by 2.3% and 2.8% for 7B and 72B models, respectively, with even greater improvements of 5.9% and 6.0% on the longest LVBench. Our code is available at https://github.com/SCZwangxiao/video-FlexReduc.git.
title AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2503.12559