How Important are Videos for Training Video LLMs?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lydakis, George, Hermans, Alexander, Athar, Ali, de Geus, Daan, Leibe, Bastian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910994973851648
author Lydakis, George
Hermans, Alexander
Athar, Ali
de Geus, Daan
Leibe, Bastian
author_facet Lydakis, George
Hermans, Alexander
Athar, Ali
de Geus, Daan
Leibe, Bastian
contents Research into Video Large Language Models (LLMs) has progressed rapidly, with numerous models and benchmarks emerging in just a few years. Typically, these models are initialized with a pretrained text-only LLM and finetuned on both image- and video-caption datasets. In this paper, we present findings indicating that Video LLMs are more capable of temporal reasoning after image-only training than one would assume, and that improvements from video-specific training are surprisingly small. Specifically, we show that image-trained versions of two LLMs trained with the recent LongVU algorithm perform significantly above chance level on TVBench, a temporal reasoning benchmark. Additionally, we introduce a simple finetuning scheme involving sequences of annotated images and questions targeting temporal capabilities. This baseline results in temporal reasoning performance close to, and occasionally higher than, what is achieved by video-trained LLMs. This suggests suboptimal utilization of rich temporal features found in real video by current models. Our analysis motivates further research into the mechanisms that allow image-trained LLMs to perform temporal reasoning, as well as into the bottlenecks that render current video training schemes inefficient.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Important are Videos for Training Video LLMs?
Lydakis, George
Hermans, Alexander
Athar, Ali
de Geus, Daan
Leibe, Bastian
Computer Vision and Pattern Recognition
Research into Video Large Language Models (LLMs) has progressed rapidly, with numerous models and benchmarks emerging in just a few years. Typically, these models are initialized with a pretrained text-only LLM and finetuned on both image- and video-caption datasets. In this paper, we present findings indicating that Video LLMs are more capable of temporal reasoning after image-only training than one would assume, and that improvements from video-specific training are surprisingly small. Specifically, we show that image-trained versions of two LLMs trained with the recent LongVU algorithm perform significantly above chance level on TVBench, a temporal reasoning benchmark. Additionally, we introduce a simple finetuning scheme involving sequences of annotated images and questions targeting temporal capabilities. This baseline results in temporal reasoning performance close to, and occasionally higher than, what is achieved by video-trained LLMs. This suggests suboptimal utilization of rich temporal features found in real video by current models. Our analysis motivates further research into the mechanisms that allow image-trained LLMs to perform temporal reasoning, as well as into the bottlenecks that render current video training schemes inefficient.
title How Important are Videos for Training Video LLMs?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06928