MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Peiran, Yu, Zhuorui, Liu, Yunze, Wu, Chi-Hao, Zhou, Enmin, Shen, Junxiao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911433176907776
author Wu, Peiran
Yu, Zhuorui
Liu, Yunze
Wu, Chi-Hao
Zhou, Enmin
Shen, Junxiao
author_facet Wu, Peiran
Yu, Zhuorui
Liu, Yunze
Wu, Chi-Hao
Zhou, Enmin
Shen, Junxiao
contents The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when extended from images to videos due to high frame rates and long durations. Token compression is a promising solution, yet most existing training-free methods cause information loss and performance degradation. To overcome this, we propose \textbf{Memory-Augmented Reinforcement Learning-based Token Compression (MARC)}, which integrates structured retrieval and RL-based distillation. MARC adopts a \textit{retrieve-then-compress} strategy using a \textbf{Visual Memory Retriever (VMR)} to select key clips and a \textbf{Compression Group Relative Policy Optimization (C-GRPO)} framework to distil reasoning ability from a teacher to a student model. Experiments on six video benchmarks show that MARC achieves near-baseline accuracy using only one frame's tokens -- reducing visual tokens by \textbf{95\%}, GPU memory by \textbf{72\%}, and latency by \textbf{23.9\%}. This demonstrates its potential for efficient, real-time video understanding in resource-constrained settings such as video QA, surveillance, and autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
Wu, Peiran
Yu, Zhuorui
Liu, Yunze
Wu, Chi-Hao
Zhou, Enmin
Shen, Junxiao
Computer Vision and Pattern Recognition
The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when extended from images to videos due to high frame rates and long durations. Token compression is a promising solution, yet most existing training-free methods cause information loss and performance degradation. To overcome this, we propose \textbf{Memory-Augmented Reinforcement Learning-based Token Compression (MARC)}, which integrates structured retrieval and RL-based distillation. MARC adopts a \textit{retrieve-then-compress} strategy using a \textbf{Visual Memory Retriever (VMR)} to select key clips and a \textbf{Compression Group Relative Policy Optimization (C-GRPO)} framework to distil reasoning ability from a teacher to a student model. Experiments on six video benchmarks show that MARC achieves near-baseline accuracy using only one frame's tokens -- reducing visual tokens by \textbf{95\%}, GPU memory by \textbf{72\%}, and latency by \textbf{23.9\%}. This demonstrates its potential for efficient, real-time video understanding in resource-constrained settings such as video QA, surveillance, and autonomous driving.
title MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.07915