Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Xiangrui, Shu, Yan, Liu, Zheng, Li, Ao, Tian, Yang, Zhao, Bo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912348710633472
author Liu, Xiangrui
Shu, Yan
Liu, Zheng
Li, Ao
Tian, Yang
Zhao, Bo
author_facet Liu, Xiangrui
Shu, Yan
Liu, Zheng
Li, Ao
Tian, Yang
Zhao, Bo
contents Despite advanced token compression techniques, existing multimodal large language models (MLLMs) still struggle with hour-long video understanding. In this work, we propose Video-XL-Pro, an efficient method for extremely long video understanding, built upon Reconstructive Compression of Tokens (ReCoT), a learnable module that leverages self-supervised learning to generate comprehensive and compact video tokens. ReCoT introduces two key components: (i) Dynamic Token Synthesizer (DTS): DTS generates pseudo-video tokens from static image tokens by learning intra-token relationships, which are then used in masked video modeling. (ii) Semantic-Guided Masking (SGM): SGM adaptively masks redundant visual tokens to facilitate more effective reconstructive learning. To improve training efficiency in MLLMs fine-tuning, we introduce a video-specific dataset pruning strategy and design a simple yet Query-aware Selector that enables the model to precisely locate query-relevant video tokens. With only 3B parameters, Video-XL-Pro outperforms most 7B models trained on larger datasets across multiple long video understanding benchmarks. Moreover, it can process over 8K frames on a single A100 GPU while maintaining high-quality performance.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Liu, Xiangrui
Shu, Yan
Liu, Zheng
Li, Ao
Tian, Yang
Zhao, Bo
Computer Vision and Pattern Recognition
Despite advanced token compression techniques, existing multimodal large language models (MLLMs) still struggle with hour-long video understanding. In this work, we propose Video-XL-Pro, an efficient method for extremely long video understanding, built upon Reconstructive Compression of Tokens (ReCoT), a learnable module that leverages self-supervised learning to generate comprehensive and compact video tokens. ReCoT introduces two key components: (i) Dynamic Token Synthesizer (DTS): DTS generates pseudo-video tokens from static image tokens by learning intra-token relationships, which are then used in masked video modeling. (ii) Semantic-Guided Masking (SGM): SGM adaptively masks redundant visual tokens to facilitate more effective reconstructive learning. To improve training efficiency in MLLMs fine-tuning, we introduce a video-specific dataset pruning strategy and design a simple yet Query-aware Selector that enables the model to precisely locate query-relevant video tokens. With only 3B parameters, Video-XL-Pro outperforms most 7B models trained on larger datasets across multiple long video understanding benchmarks. Moreover, it can process over 8K frames on a single A100 GPU while maintaining high-quality performance.
title Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18478