Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zeqian, Di, Shangzhe, Zhai, Zhonghua, Huang, Weilin, Wang, Yanfeng, Xie, Weidi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918212177756160
author Li, Zeqian
Di, Shangzhe
Zhai, Zhonghua
Huang, Weilin
Wang, Yanfeng
Xie, Weidi
author_facet Li, Zeqian
Di, Shangzhe
Zhai, Zhonghua
Huang, Weilin
Wang, Yanfeng
Xie, Weidi
contents This paper presents a computational model for universal video temporal grounding, which accurately localizes temporal moments in videos based on natural language queries (e.g., questions or descriptions). Unlike existing methods that are often limited to specific video domains or durations, we propose UniTime, a robust and universal video grounding model leveraging the strong vision-language understanding capabilities of generative Multi-modal Large Language Models (MLLMs). Our model effectively handles videos of diverse views, genres, and lengths while comprehending complex language queries. The key contributions include: (i) We consider steering strong MLLMs for temporal grounding in videos. To enable precise timestamp outputs, we incorporate temporal information by interleaving timestamp tokens with video tokens. (ii) By training the model to handle videos with different input granularities through adaptive frame scaling, our approach achieves robust temporal grounding for both short and long videos. (iii) Comprehensive experiments show that UniTime outperforms state-of-the-art approaches in both zero-shot and dataset-specific finetuned settings across five public temporal grounding benchmarks. (iv) When employed as a preliminary moment retriever for long-form video question-answering (VideoQA), UniTime significantly improves VideoQA accuracy, highlighting its value for complex video understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18883
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
Li, Zeqian
Di, Shangzhe
Zhai, Zhonghua
Huang, Weilin
Wang, Yanfeng
Xie, Weidi
Computer Vision and Pattern Recognition
This paper presents a computational model for universal video temporal grounding, which accurately localizes temporal moments in videos based on natural language queries (e.g., questions or descriptions). Unlike existing methods that are often limited to specific video domains or durations, we propose UniTime, a robust and universal video grounding model leveraging the strong vision-language understanding capabilities of generative Multi-modal Large Language Models (MLLMs). Our model effectively handles videos of diverse views, genres, and lengths while comprehending complex language queries. The key contributions include: (i) We consider steering strong MLLMs for temporal grounding in videos. To enable precise timestamp outputs, we incorporate temporal information by interleaving timestamp tokens with video tokens. (ii) By training the model to handle videos with different input granularities through adaptive frame scaling, our approach achieves robust temporal grounding for both short and long videos. (iii) Comprehensive experiments show that UniTime outperforms state-of-the-art approaches in both zero-shot and dataset-specific finetuned settings across five public temporal grounding benchmarks. (iv) When employed as a preliminary moment retriever for long-form video question-answering (VideoQA), UniTime significantly improves VideoQA accuracy, highlighting its value for complex video understanding tasks.
title Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18883