TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yunxiao, Liu, Meng, Liu, Wenqi, Song, Xuemeng, Wen, Bin, Yang, Fan, Gao, Tingting, Zhang, Di, Zhou, Guorui, Nie, Liqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912525740670976
author Wang, Yunxiao
Liu, Meng
Liu, Wenqi
Song, Xuemeng
Wen, Bin
Yang, Fan
Gao, Tingting
Zhang, Di
Zhou, Guorui
Nie, Liqiang
author_facet Wang, Yunxiao
Liu, Meng
Liu, Wenqi
Song, Xuemeng
Wen, Bin
Yang, Fan
Gao, Tingting
Zhang, Di
Zhou, Guorui
Nie, Liqiang
contents Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning dataset that focuses on enhancing temporal comprehension across five key dimensions. In order to reduce reliance on costly temporal annotations, we introduce a multi-task prompt fine-tuning approach that seamlessly integrates temporal-sensitive tasks into existing instruction datasets without requiring additional annotations. Furthermore, we develop a novel benchmark for temporal-sensitive video understanding that not only fills the gaps in dimension coverage left by existing benchmarks but also rigorously filters out potential shortcuts, ensuring a more accurate evaluation. Extensive experimental results demonstrate that our approach significantly enhances the temporal understanding of video-LLMs while avoiding reliance on shortcuts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
Wang, Yunxiao
Liu, Meng
Liu, Wenqi
Song, Xuemeng
Wen, Bin
Yang, Fan
Gao, Tingting
Zhang, Di
Zhou, Guorui
Nie, Liqiang
Computer Vision and Pattern Recognition
Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning dataset that focuses on enhancing temporal comprehension across five key dimensions. In order to reduce reliance on costly temporal annotations, we introduce a multi-task prompt fine-tuning approach that seamlessly integrates temporal-sensitive tasks into existing instruction datasets without requiring additional annotations. Furthermore, we develop a novel benchmark for temporal-sensitive video understanding that not only fills the gaps in dimension coverage left by existing benchmarks but also rigorously filters out potential shortcuts, ensuring a more accurate evaluation. Extensive experimental results demonstrate that our approach significantly enhances the temporal understanding of video-LLMs while avoiding reliance on shortcuts.
title TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.09994