Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Woo, Jongbhin, Ryu, Hyeonggon, Jang, Youngjoon, Cho, Jae Won, Chung, Joon Son
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914976671727616
author Woo, Jongbhin
Ryu, Hyeonggon
Jang, Youngjoon
Cho, Jae Won
Chung, Joon Son
author_facet Woo, Jongbhin
Ryu, Hyeonggon
Jang, Youngjoon
Cho, Jae Won
Chung, Joon Son
contents Video Temporal Grounding (VTG) aims to identify visual frames in a video clip that match text queries. Recent studies in VTG employ cross-attention to correlate visual frames and text queries as individual token sequences. However, these approaches overlook a crucial aspect of the problem: a holistic understanding of the query sentence. A model may capture correlations between individual word tokens and arbitrary visual frames while possibly missing out on the global meaning. To address this, we introduce two primary contributions: (1) a visual frame-level gate mechanism that incorporates holistic textual information, (2) cross-modal alignment loss to learn the fine-grained correlation between query and relevant frames. As a result, we regularize the effect of individual word tokens and suppress irrelevant visual frames. We demonstrate that our method outperforms state-of-the-art approaches in VTG benchmarks, indicating that holistic text understanding guides the model to focus on the semantically important parts within the video.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13598
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
Woo, Jongbhin
Ryu, Hyeonggon
Jang, Youngjoon
Cho, Jae Won
Chung, Joon Son
Computer Vision and Pattern Recognition
Video Temporal Grounding (VTG) aims to identify visual frames in a video clip that match text queries. Recent studies in VTG employ cross-attention to correlate visual frames and text queries as individual token sequences. However, these approaches overlook a crucial aspect of the problem: a holistic understanding of the query sentence. A model may capture correlations between individual word tokens and arbitrary visual frames while possibly missing out on the global meaning. To address this, we introduce two primary contributions: (1) a visual frame-level gate mechanism that incorporates holistic textual information, (2) cross-modal alignment loss to learn the fine-grained correlation between query and relevant frames. As a result, we regularize the effect of individual word tokens and suppress irrelevant visual frames. We demonstrate that our method outperforms state-of-the-art approaches in VTG benchmarks, indicating that holistic text understanding guides the model to focus on the semantically important parts within the video.
title Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.13598