Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Qian, Su, Lixin, Zhao, Jiashu, Xia, Long, Cai, Hengyi, Cheng, Suqi, Tang, Hengzhu, Wang, Junfeng, Yin, Dawei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910288243064832
author Li, Qian
Su, Lixin
Zhao, Jiashu
Xia, Long
Cai, Hengyi
Cheng, Suqi
Tang, Hengzhu
Wang, Junfeng
Yin, Dawei
author_facet Li, Qian
Su, Lixin
Zhao, Jiashu
Xia, Long
Cai, Hengyi
Cheng, Suqi
Tang, Hengzhu
Wang, Junfeng
Yin, Dawei
contents Text-video retrieval is a challenging task that aims to identify relevant videos given textual queries. Compared to conventional textual retrieval, the main obstacle for text-video retrieval is the semantic gap between the textual nature of queries and the visual richness of video content. Previous works primarily focus on aligning the query and the video by finely aggregating word-frame matching signals. Inspired by the human cognitive process of modularly judging the relevance between text and video, the judgment needs high-order matching signal due to the consecutive and complex nature of video contents. In this paper, we propose chunk-level text-video matching, where the query chunks are extracted to describe a specific retrieval unit, and the video chunks are segmented into distinct clips from videos. We formulate the chunk-level matching as n-ary correlations modeling between words of the query and frames of the video and introduce a multi-modal hypergraph for n-ary correlation modeling. By representing textual units and video frames as nodes and using hyperedges to depict their relationships, a multi-modal hypergraph is constructed. In this way, the query and the video can be aligned in a high-order semantic space. In addition, to enhance the model's generalization ability, the extracted features are fed into a variational inference component for computation, obtaining the variational representation under the Gaussian distribution. The incorporation of hypergraphs and variational inference allows our model to capture complex, n-ary interactions among textual and visual contents. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on the text-video retrieval task.
format Preprint
id arxiv_https___arxiv_org_abs_2401_03177
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
Li, Qian
Su, Lixin
Zhao, Jiashu
Xia, Long
Cai, Hengyi
Cheng, Suqi
Tang, Hengzhu
Wang, Junfeng
Yin, Dawei
Computer Vision and Pattern Recognition
Computation and Language
Text-video retrieval is a challenging task that aims to identify relevant videos given textual queries. Compared to conventional textual retrieval, the main obstacle for text-video retrieval is the semantic gap between the textual nature of queries and the visual richness of video content. Previous works primarily focus on aligning the query and the video by finely aggregating word-frame matching signals. Inspired by the human cognitive process of modularly judging the relevance between text and video, the judgment needs high-order matching signal due to the consecutive and complex nature of video contents. In this paper, we propose chunk-level text-video matching, where the query chunks are extracted to describe a specific retrieval unit, and the video chunks are segmented into distinct clips from videos. We formulate the chunk-level matching as n-ary correlations modeling between words of the query and frames of the video and introduce a multi-modal hypergraph for n-ary correlation modeling. By representing textual units and video frames as nodes and using hyperedges to depict their relationships, a multi-modal hypergraph is constructed. In this way, the query and the video can be aligned in a high-order semantic space. In addition, to enhance the model's generalization ability, the extracted features are fed into a variational inference component for computation, obtaining the variational representation under the Gaussian distribution. The incorporation of hypergraphs and variational inference allows our model to capture complex, n-ary interactions among textual and visual contents. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on the text-video retrieval task.
title Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2401.03177