Unleashing Hour-Scale Video Training for Long Video-Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Jingyang, Wu, Jialian, Sun, Ximeng, Wang, Ze, Liu, Jiang, Su, Yusheng, Yu, Xiaodong, Chen, Hao, Luo, Jiebo, Liu, Zicheng, Barsoum, Emad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908685602652160
author Lin, Jingyang
Wu, Jialian
Sun, Ximeng
Wang, Ze
Liu, Jiang
Su, Yusheng
Yu, Xiaodong
Chen, Hao
Luo, Jiebo
Liu, Zicheng
Barsoum, Emad
author_facet Lin, Jingyang
Wu, Jialian
Sun, Ximeng
Wang, Ze
Liu, Jiang
Su, Yusheng
Yu, Xiaodong
Chen, Hao
Luo, Jiebo
Liu, Zicheng
Barsoum, Emad
contents Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unleashing Hour-Scale Video Training for Long Video-Language Understanding
Lin, Jingyang
Wu, Jialian
Sun, Ximeng
Wang, Ze
Liu, Jiang
Su, Yusheng
Yu, Xiaodong
Chen, Hao
Luo, Jiebo
Liu, Zicheng
Barsoum, Emad
Computer Vision and Pattern Recognition
Computation and Language
Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.
title Unleashing Hour-Scale Video Training for Long Video-Language Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.05332