ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiao, Si, Qingyi, Wu, Jianlong, Zhu, Shiyu, Cao, Li, Nie, Liqiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910889318285312
author Wang, Xiao
Si, Qingyi
Wu, Jianlong
Zhu, Shiyu
Cao, Li
Nie, Liqiang
author_facet Wang, Xiao
Si, Qingyi
Wu, Jianlong
Zhu, Shiyu
Cao, Li
Nie, Liqiang
contents Video Large Language Models (VideoLLMs) have made significant strides in video understanding but struggle with long videos due to the limitations of their backbone LLMs. Existing solutions rely on length extrapolation, which is memory-constrained, or visual token compression, which primarily leverages low-level temporal redundancy while overlooking the more effective high-level knowledge redundancy. To address this, we propose $\textbf{ReTaKe}$, a training-free method with two novel modules DPSelect and PivotKV, to jointly reduce both temporal visual redundancy and knowledge redundancy for video compression. To align with the way of human temporal perception, DPSelect identifies keyframes based on inter-frame distance peaks. To leverage LLMs' learned prior knowledge, PivotKV marks the keyframes as pivots and compress non-pivot frames by pruning low-attention tokens in their KV cache. ReTaKe enables VideoLLMs to process 8 times longer frames (up to 2048), outperforming similar-sized models by 3-5% and even rivaling much larger ones on VideoMME, MLVU, LongVideoBench, and LVBench. Moreover, by overlapping compression operations with prefilling, ReTaKe introduces only ~10% prefilling latency overhead while reducing decoding latency by ~20%. Our code is available at https://github.com/SCZwangxiao/video-ReTaKe.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20504
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
Wang, Xiao
Si, Qingyi
Wu, Jianlong
Zhu, Shiyu
Cao, Li
Nie, Liqiang
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Video Large Language Models (VideoLLMs) have made significant strides in video understanding but struggle with long videos due to the limitations of their backbone LLMs. Existing solutions rely on length extrapolation, which is memory-constrained, or visual token compression, which primarily leverages low-level temporal redundancy while overlooking the more effective high-level knowledge redundancy. To address this, we propose $\textbf{ReTaKe}$, a training-free method with two novel modules DPSelect and PivotKV, to jointly reduce both temporal visual redundancy and knowledge redundancy for video compression. To align with the way of human temporal perception, DPSelect identifies keyframes based on inter-frame distance peaks. To leverage LLMs' learned prior knowledge, PivotKV marks the keyframes as pivots and compress non-pivot frames by pruning low-attention tokens in their KV cache. ReTaKe enables VideoLLMs to process 8 times longer frames (up to 2048), outperforming similar-sized models by 3-5% and even rivaling much larger ones on VideoMME, MLVU, LongVideoBench, and LVBench. Moreover, by overlapping compression operations with prefilling, ReTaKe introduces only ~10% prefilling latency overhead while reducing decoding latency by ~20%. Our code is available at https://github.com/SCZwangxiao/video-ReTaKe.
title ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2412.20504