Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jeong, Boseung, Park, Jicheol, Kim, Sungyeon, Kwak, Suha
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915225425412096
author Jeong, Boseung
Park, Jicheol
Kim, Sungyeon
Kwak, Suha
author_facet Jeong, Boseung
Park, Jicheol
Kim, Sungyeon
Kwak, Suha
contents Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and textual features and often ignore audio, although it helps enhance overall comprehension of video content. Moreover, traditional models that incorporate audio blindly utilize the audio input regardless of whether it is useful or not, resulting in suboptimal video representation. To address these limitations, we propose a novel video-text retrieval framework, Audio-guided VIdeo representation learning with GATEd attention (AVIGATE), that effectively leverages audio cues through a gated attention mechanism that selectively filters out uninformative audio signals. In addition, we propose an adaptive margin-based contrastive loss to deal with the inherently unclear positive-negative relationship between video and text, which facilitates learning better video-text alignment. Our extensive experiments demonstrate that AVIGATE achieves state-of-the-art performance on all the public benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02397
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
Jeong, Boseung
Park, Jicheol
Kim, Sungyeon
Kwak, Suha
Computer Vision and Pattern Recognition
Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and textual features and often ignore audio, although it helps enhance overall comprehension of video content. Moreover, traditional models that incorporate audio blindly utilize the audio input regardless of whether it is useful or not, resulting in suboptimal video representation. To address these limitations, we propose a novel video-text retrieval framework, Audio-guided VIdeo representation learning with GATEd attention (AVIGATE), that effectively leverages audio cues through a gated attention mechanism that selectively filters out uninformative audio signals. In addition, we propose an adaptive margin-based contrastive loss to deal with the inherently unclear positive-negative relationship between video and text, which facilitates learning better video-text alignment. Our extensive experiments demonstrate that AVIGATE achieves state-of-the-art performance on all the public benchmarks.
title Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.02397