Multimodal Long Video Modeling Based on Temporal Dynamic Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hao, Haoran, Han, Jiaming, Zhang, Yiyuan, Yue, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912326792249344
author Hao, Haoran
Han, Jiaming
Zhang, Yiyuan
Yue, Xiangyu
author_facet Hao, Haoran
Han, Jiaming
Zhang, Yiyuan
Yue, Xiangyu
contents Recent advances in Large Language Models (LLMs) have led to significant breakthroughs in video understanding. However, existing models still struggle with long video processing due to the context length constraint of LLMs and the vast amount of information within the video. Although some recent methods are designed for long video understanding, they often lose crucial information during token compression and struggle with additional modality like audio. In this work, we propose a dynamic long video encoding method utilizing the temporal relationship between frames, named Temporal Dynamic Context (TDC). Firstly, we segment the video into semantically consistent scenes based on inter-frame similarities, then encode each frame into tokens using visual-audio encoders. Secondly, we propose a novel temporal context compressor to reduce the number of tokens within each segment. Specifically, we employ a query-based Transformer to aggregate video, audio, and instruction text tokens into a limited set of temporal context tokens. Finally, we feed the static frame tokens and the temporal context tokens into the LLM for video understanding. Furthermore, to handle extremely long videos, we propose a training-free chain-of-thought strategy that progressively extracts answers from multiple video segments. These intermediate answers serve as part of the reasoning process and contribute to the final answer. We conduct extensive experiments on general video understanding and audio-video understanding benchmarks, where our method demonstrates strong performance. The code and models are available at https://github.com/Hoar012/TDC-Video.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10443
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Long Video Modeling Based on Temporal Dynamic Context
Hao, Haoran
Han, Jiaming
Zhang, Yiyuan
Yue, Xiangyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
Recent advances in Large Language Models (LLMs) have led to significant breakthroughs in video understanding. However, existing models still struggle with long video processing due to the context length constraint of LLMs and the vast amount of information within the video. Although some recent methods are designed for long video understanding, they often lose crucial information during token compression and struggle with additional modality like audio. In this work, we propose a dynamic long video encoding method utilizing the temporal relationship between frames, named Temporal Dynamic Context (TDC). Firstly, we segment the video into semantically consistent scenes based on inter-frame similarities, then encode each frame into tokens using visual-audio encoders. Secondly, we propose a novel temporal context compressor to reduce the number of tokens within each segment. Specifically, we employ a query-based Transformer to aggregate video, audio, and instruction text tokens into a limited set of temporal context tokens. Finally, we feed the static frame tokens and the temporal context tokens into the LLM for video understanding. Furthermore, to handle extremely long videos, we propose a training-free chain-of-thought strategy that progressively extracts answers from multiple video segments. These intermediate answers serve as part of the reasoning process and contribute to the final answer. We conduct extensive experiments on general video understanding and audio-video understanding benchmarks, where our method demonstrates strong performance. The code and models are available at https://github.com/Hoar012/TDC-Video.
title Multimodal Long Video Modeling Based on Temporal Dynamic Context
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
url https://arxiv.org/abs/2504.10443