Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Yuxiao, Wang, Jue, Zhang, Zhikang, Yi, Jingru, Zhang, Xu, Zou, Yang, Cai, Zhaowei, Yuan, Jianbo, Li, Xinyu, Yang, Hao, Modolo, Davide
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917283823091712
author Chen, Yuxiao
Wang, Jue
Zhang, Zhikang
Yi, Jingru
Zhang, Xu
Zou, Yang
Cai, Zhaowei
Yuan, Jianbo
Li, Xinyu
Yang, Hao
Modolo, Davide
author_facet Chen, Yuxiao
Wang, Jue
Zhang, Zhikang
Yi, Jingru
Zhang, Xu
Zou, Yang
Cai, Zhaowei
Yuan, Jianbo
Li, Xinyu
Yang, Hao
Modolo, Davide
contents With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens of minutes has become both feasible and increasingly prevalent. However, the inherently redundant nature of video sequences poses significant challenges for contemporary state-of-the-art models. These challenges stem from two primary aspects: 1) efficiently incorporating a larger number of frames within memory constraints, and 2) extracting discriminative information from the vast volume of input data. In this paper, we introduce a novel end-to-end schema for long-form video understanding, which includes an information-density-based adaptive video sampler (AVS) and an autoencoder-based spatiotemporal video compressor (SVC) integrated with a multimodal large language model (MLLM). Our proposed system offers two major advantages: it adaptively and effectively captures essential information from video sequences of varying durations, and it achieves high compression rates while preserving crucial discriminative information. The proposed framework demonstrates promising performance across various benchmarks, excelling in both long-form video understanding tasks and standard video understanding benchmarks. These results underscore the versatility and efficacy of our approach, particularly in managing the complexities of prolonged video sequences.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17869
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
Chen, Yuxiao
Wang, Jue
Zhang, Zhikang
Yi, Jingru
Zhang, Xu
Zou, Yang
Cai, Zhaowei
Yuan, Jianbo
Li, Xinyu
Yang, Hao
Modolo, Davide
Computer Vision and Pattern Recognition
With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens of minutes has become both feasible and increasingly prevalent. However, the inherently redundant nature of video sequences poses significant challenges for contemporary state-of-the-art models. These challenges stem from two primary aspects: 1) efficiently incorporating a larger number of frames within memory constraints, and 2) extracting discriminative information from the vast volume of input data. In this paper, we introduce a novel end-to-end schema for long-form video understanding, which includes an information-density-based adaptive video sampler (AVS) and an autoencoder-based spatiotemporal video compressor (SVC) integrated with a multimodal large language model (MLLM). Our proposed system offers two major advantages: it adaptively and effectively captures essential information from video sequences of varying durations, and it achieves high compression rates while preserving crucial discriminative information. The proposed framework demonstrates promising performance across various benchmarks, excelling in both long-form video understanding tasks and standard video understanding benchmarks. These results underscore the versatility and efficacy of our approach, particularly in managing the complexities of prolonged video sequences.
title Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.17869