Video sentence grounding with temporally global textual knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Cai, Zhang, Runzhong, Gao, Jianjun, Wu, Kejun, Yap, Kim-Hui, Wang, Yi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916268837175296
author Chen, Cai
Zhang, Runzhong
Gao, Jianjun
Wu, Kejun
Yap, Kim-Hui
Wang, Yi
author_facet Chen, Cai
Zhang, Runzhong
Gao, Jianjun
Wu, Kejun
Yap, Kim-Hui
Wang, Yi
contents Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent domain gap between different modalities. In this paper, we utilize pseudo-query features containing extensive temporally global textual knowledge sourced from the same video-query pair, to enhance the bridging of domain gaps and attain a heightened level of similarity between multi-modal features. Specifically, we propose a Pseudo-query Intermediary Network (PIN) to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning. Subsequently, we utilize learnable prompts to encapsulate the knowledge of pseudo-queries, propagating them into the textual encoder and multi-modal fusion module, further enhancing the feature alignment between visual and language for better temporal grounding. Extensive experiments conducted on the Charades-STA and ActivityNet-Captions datasets demonstrate the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2404_13611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video sentence grounding with temporally global textual knowledge
Chen, Cai
Zhang, Runzhong
Gao, Jianjun
Wu, Kejun
Yap, Kim-Hui
Wang, Yi
Computer Vision and Pattern Recognition
Computation and Language
Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent domain gap between different modalities. In this paper, we utilize pseudo-query features containing extensive temporally global textual knowledge sourced from the same video-query pair, to enhance the bridging of domain gaps and attain a heightened level of similarity between multi-modal features. Specifically, we propose a Pseudo-query Intermediary Network (PIN) to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning. Subsequently, we utilize learnable prompts to encapsulate the knowledge of pseudo-queries, propagating them into the textual encoder and multi-modal fusion module, further enhancing the feature alignment between visual and language for better temporal grounding. Extensive experiments conducted on the Charades-STA and ActivityNet-Captions datasets demonstrate the effectiveness of our method.
title Video sentence grounding with temporally global textual knowledge
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2404.13611