An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qi, Ji, Yao, Yuan, Bai, Yushi, Xu, Bin, Li, Juanzi, Liu, Zhiyuan, Chua, Tat-Seng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908329635217408
author Qi, Ji
Yao, Yuan
Bai, Yushi
Xu, Bin
Li, Juanzi
Liu, Zhiyuan
Chua, Tat-Seng
author_facet Qi, Ji
Yao, Yuan
Bai, Yushi
Xu, Bin
Li, Juanzi
Liu, Zhiyuan
Chua, Tat-Seng
contents Large Multimodal Models (LMMs) uniformly perceive video frames, creating computational inefficiency for videos with inherently varying temporal information density. This paper present \textbf{Quicksviewer}, an LMM with new perceiving paradigm that partitions a video of nonuniform density into varying cubes using Gumbel Softmax, followed by a unified resampling for each cube to achieve efficient video understanding. This simple and intuitive approach dynamically compress video online based on its temporal density, significantly reducing spatiotemporal redundancy (overall 45$\times$ compression rate), while enabling efficient training with large receptive field. We train the model from a language backbone through three progressive stages, each incorporating lengthy videos on average of 420s/1fps thanks to the perceiving efficiency. With only 0.8M total video-text samples for training, our model outperforms the direct baseline employing a fixed partitioning strategy by a maximum of 8.72 in accuracy, demonstrating the effectiveness in performance. On Video-MME, Quicksviewer achieves SOTA under modest sequence lengths using just up to 5\% of tokens per frame required by baselines. With this paradigm, scaling up the number of input frames reveals a clear power law of the model capabilities. It is also empirically verified that the segments generated by the cubing network can help for analyzing continuous events in videos.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
Qi, Ji
Yao, Yuan
Bai, Yushi
Xu, Bin
Li, Juanzi
Liu, Zhiyuan
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Computation and Language
Large Multimodal Models (LMMs) uniformly perceive video frames, creating computational inefficiency for videos with inherently varying temporal information density. This paper present \textbf{Quicksviewer}, an LMM with new perceiving paradigm that partitions a video of nonuniform density into varying cubes using Gumbel Softmax, followed by a unified resampling for each cube to achieve efficient video understanding. This simple and intuitive approach dynamically compress video online based on its temporal density, significantly reducing spatiotemporal redundancy (overall 45$\times$ compression rate), while enabling efficient training with large receptive field. We train the model from a language backbone through three progressive stages, each incorporating lengthy videos on average of 420s/1fps thanks to the perceiving efficiency. With only 0.8M total video-text samples for training, our model outperforms the direct baseline employing a fixed partitioning strategy by a maximum of 8.72 in accuracy, demonstrating the effectiveness in performance. On Video-MME, Quicksviewer achieves SOTA under modest sequence lengths using just up to 5\% of tokens per frame required by baselines. With this paradigm, scaling up the number of input frames reveals a clear power law of the model capabilities. It is also empirically verified that the segments generated by the cubing network can help for analyzing continuous events in videos.
title An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2504.15270