TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Wei-Yuan, Chang, Kai-Po, Huang, Chi-Pin, Yang, Fu-En, Wang, Yu-Chiang Frank
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908749901332480
author Cheng, Wei-Yuan
Chang, Kai-Po
Huang, Chi-Pin
Yang, Fu-En
Wang, Yu-Chiang Frank
author_facet Cheng, Wei-Yuan
Chang, Kai-Po
Huang, Chi-Pin
Yang, Fu-En
Wang, Yu-Chiang Frank
contents Dense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data. However, existing VideoLLMs remain challenging in identifying precise event boundaries in untrimmed videos, causing the generated captions to be not properly grounded. In this paper, we propose TA-Prompting, which enhances VideoLLMs via Temporal Anchors that learn to precisely localize events and prompt the VideoLLMs to perform temporal-aware video event understanding. During inference, in order to properly determine the output caption sequence from an arbitrary number of events presented within a video, we introduce an event coherent sampling strategy to select event captions with sufficient coherence across temporal events and cross-modal similarity with the given video. Through extensive experiments on benchmark datasets, we show that our TA-Prompting is favorable against state-of-the-art VideoLLMs, yielding superior performance on dense video captioning and temporal understanding tasks including moment retrieval and temporalQA.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02908
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
Cheng, Wei-Yuan
Chang, Kai-Po
Huang, Chi-Pin
Yang, Fu-En
Wang, Yu-Chiang Frank
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Dense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data. However, existing VideoLLMs remain challenging in identifying precise event boundaries in untrimmed videos, causing the generated captions to be not properly grounded. In this paper, we propose TA-Prompting, which enhances VideoLLMs via Temporal Anchors that learn to precisely localize events and prompt the VideoLLMs to perform temporal-aware video event understanding. During inference, in order to properly determine the output caption sequence from an arbitrary number of events presented within a video, we introduce an event coherent sampling strategy to select event captions with sufficient coherence across temporal events and cross-modal similarity with the given video. Through extensive experiments on benchmark datasets, we show that our TA-Prompting is favorable against state-of-the-art VideoLLMs, yielding superior performance on dense video captioning and temporal understanding tasks including moment retrieval and temporalQA.
title TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.02908