ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ma, David, Yuan, Huaqing, Wang, Xingjian, Zang, Qianbo, Liu, Tianci, He, Xinyang, Wei, Yanbin, Guo, Jiawei, Jiahui, Ni, Yang, Zhenzhu, Cao, Meng, Quan, Shanghaoran, Li, Yizhi, Zhou, Wangchunshu, Liu, Jiaheng, Huang, Wenhao, Zhang, Ge, Ni, Shiwen, Jin, Xiaojie
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918038906863616
author Ma, David
Yuan, Huaqing
Wang, Xingjian
Zang, Qianbo
Liu, Tianci
He, Xinyang
Wei, Yanbin
Guo, Jiawei
Jiahui, Ni
Yang, Zhenzhu
Cao, Meng
Quan, Shanghaoran
Li, Yizhi
Zhou, Wangchunshu
Liu, Jiaheng
Huang, Wenhao
Zhang, Ge
Ni, Shiwen
Jin, Xiaojie
author_facet Ma, David
Yuan, Huaqing
Wang, Xingjian
Zang, Qianbo
Liu, Tianci
He, Xinyang
Wei, Yanbin
Guo, Jiawei
Jiahui, Ni
Yang, Zhenzhu
Cao, Meng
Quan, Shanghaoran
Li, Yizhi
Zhou, Wangchunshu
Liu, Jiaheng
Huang, Wenhao
Zhang, Ge
Ni, Shiwen
Jin, Xiaojie
contents Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect this multi-scale design or scatter scale-specific questions across different videos, preventing direct comparison of model performance across timescales on the same content. To address this, we introduce ScaleLong, the first benchmark to disentangle these factors by embedding questions targeting four hierarchical timescales -- clip (seconds), shot (tens of seconds), event (minutes), and story (hours) -- all within the same video content. This within-content multi-timescale questioning design enables direct comparison of model performance across timescales on identical videos. ScaleLong features 269 long videos (avg.\ 86\,min) from 5 main categories and 36 sub-categories, with 4--8 carefully designed questions, including at least one question for each timescale. Evaluating 23 MLLMs reveals a U-shaped performance curve, with higher accuracy at the shortest and longest timescales and a dip at intermediate levels. Furthermore, ablation studies show that increased visual token capacity consistently enhances reasoning across all timescales. ScaleLong offers a fine-grained, multi-timescale benchmark for advancing MLLM capabilities in long-video understanding. The code and dataset are available https://github.com/multimodal-art-projection/ScaleLong.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
Ma, David
Yuan, Huaqing
Wang, Xingjian
Zang, Qianbo
Liu, Tianci
He, Xinyang
Wei, Yanbin
Guo, Jiawei
Jiahui, Ni
Yang, Zhenzhu
Cao, Meng
Quan, Shanghaoran
Li, Yizhi
Zhou, Wangchunshu
Liu, Jiaheng
Huang, Wenhao
Zhang, Ge
Ni, Shiwen
Jin, Xiaojie
Computer Vision and Pattern Recognition
Computation and Language
Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect this multi-scale design or scatter scale-specific questions across different videos, preventing direct comparison of model performance across timescales on the same content. To address this, we introduce ScaleLong, the first benchmark to disentangle these factors by embedding questions targeting four hierarchical timescales -- clip (seconds), shot (tens of seconds), event (minutes), and story (hours) -- all within the same video content. This within-content multi-timescale questioning design enables direct comparison of model performance across timescales on identical videos. ScaleLong features 269 long videos (avg.\ 86\,min) from 5 main categories and 36 sub-categories, with 4--8 carefully designed questions, including at least one question for each timescale. Evaluating 23 MLLMs reveals a U-shaped performance curve, with higher accuracy at the shortest and longest timescales and a dip at intermediate levels. Furthermore, ablation studies show that increased visual token capacity consistently enhances reasoning across all timescales. ScaleLong offers a fine-grained, multi-timescale benchmark for advancing MLLM capabilities in long-video understanding. The code and dataset are available https://github.com/multimodal-art-projection/ScaleLong.
title ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.23922