METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Mengyue, Chen, Shuo, Kersting, Kristian, Tresp, Volker, Ma, Yunpu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183306489856
author Wang, Mengyue
Chen, Shuo
Kersting, Kristian
Tresp, Volker
Ma, Yunpu
author_facet Wang, Mengyue
Chen, Shuo
Kersting, Kristian
Tresp, Volker
Ma, Yunpu
contents Recent advances in Video Large Language Models (VLLMs) have significantly enhanced their ability to understand video content. Nonetheless, processing long videos remains challenging due to high computational demands and the redundancy present in the visual data. In this work, we propose METok, a training-free, Multi-stage Event-based Token compression framework designed to accelerate VLLMs' inference while preserving accuracy. METok progressively eliminates redundant visual tokens across three critical stages: (1) event-aware compression during vision encoding, (2) hierarchical token pruning in the prefilling stage based on semantic alignment and event importance, and (3) a decoding-stage KV Cache optimization that further reduces memory consumption. Our experiments on diverse video benchmarks demonstrate that METok achieves an optimal trade-off between efficiency and accuracy by dynamically selecting informative visual tokens. For instance, equipping LongVA-7B with METok realizes an 80.6% FLOPs reduction and 93.5% KV Cache memory savings, all while maintaining comparable or even superior accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02850
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
Wang, Mengyue
Chen, Shuo
Kersting, Kristian
Tresp, Volker
Ma, Yunpu
Computer Vision and Pattern Recognition
Recent advances in Video Large Language Models (VLLMs) have significantly enhanced their ability to understand video content. Nonetheless, processing long videos remains challenging due to high computational demands and the redundancy present in the visual data. In this work, we propose METok, a training-free, Multi-stage Event-based Token compression framework designed to accelerate VLLMs' inference while preserving accuracy. METok progressively eliminates redundant visual tokens across three critical stages: (1) event-aware compression during vision encoding, (2) hierarchical token pruning in the prefilling stage based on semantic alignment and event importance, and (3) a decoding-stage KV Cache optimization that further reduces memory consumption. Our experiments on diverse video benchmarks demonstrate that METok achieves an optimal trade-off between efficiency and accuracy by dynamically selecting informative visual tokens. For instance, equipping LongVA-7B with METok realizes an 80.6% FLOPs reduction and 93.5% KV Cache memory savings, all while maintaining comparable or even superior accuracy.
title METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.02850