Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nguyen, Thong, Hu, Zhiyuan, Lin, Xu, Nguyen, Cong-Duy, Ng, See-Kiong, Tuan, Luu Anh
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910951738966016
author Nguyen, Thong
Hu, Zhiyuan
Lin, Xu
Nguyen, Cong-Duy
Ng, See-Kiong
Tuan, Luu Anh
author_facet Nguyen, Thong
Hu, Zhiyuan
Lin, Xu
Nguyen, Cong-Duy
Ng, See-Kiong
Tuan, Luu Anh
contents Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal understanding capacity. As such, they have not deciphered important components that contribute to temporal understanding ability, which might limit the potential of these LVLMs for video understanding. In this work, we conduct a thorough empirical study to demystify crucial components that influence the temporal understanding of LVLMs. Our empirical study reveals that significant impacts are centered around the intermediate interface between the visual encoder and the large language model. Building on these insights, we propose a temporal-oriented recipe that encompasses temporal-oriented training schemes and an upscaled interface. Our final model developed using our recipe significantly enhances previous LVLMs on standard video understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
Nguyen, Thong
Hu, Zhiyuan
Lin, Xu
Nguyen, Cong-Duy
Ng, See-Kiong
Tuan, Luu Anh
Computer Vision and Pattern Recognition
Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal understanding capacity. As such, they have not deciphered important components that contribute to temporal understanding ability, which might limit the potential of these LVLMs for video understanding. In this work, we conduct a thorough empirical study to demystify crucial components that influence the temporal understanding of LVLMs. Our empirical study reveals that significant impacts are centered around the intermediate interface between the visual encoder and the large language model. Building on these insights, we propose a temporal-oriented recipe that encompasses temporal-oriented training schemes and an upscaled interface. Our final model developed using our recipe significantly enhances previous LVLMs on standard video understanding tasks.
title Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.12605