Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xie, Huang, Khorrami, Khazar, Räsänen, Okko, Virtanen, Tuomas
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912002923823104
author Xie, Huang
Khorrami, Khazar
Räsänen, Okko
Virtanen, Tuomas
author_facet Xie, Huang
Khorrami, Khazar
Räsänen, Okko
Virtanen, Tuomas
contents Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from pairs of audio samples and their human-provided captions, categorizing each pair as either positive or negative. This may result in suboptimal systems due to varying levels of relevance between audio samples and captions. In contrast, a recent study used human-assigned relevance ratings, i.e., continuous relevances, for these pairs but did not obtain performance gains in audio-text relevance learning. This work introduces a relevance learning method that utilizes both human-assigned continuous relevance ratings and binary relevances using a combination of a listwise ranking objective and a contrastive learning objective. Experimental results demonstrate the effectiveness of the proposed method, showing improvements in language-based audio retrieval, a downstream task in audio-text relevance learning. In addition, we analyze how properties of the captions or audio clips contribute to the continuous audio-text relevances provided by humans or learned by the machine.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
Xie, Huang
Khorrami, Khazar
Räsänen, Okko
Virtanen, Tuomas
Audio and Speech Processing
Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from pairs of audio samples and their human-provided captions, categorizing each pair as either positive or negative. This may result in suboptimal systems due to varying levels of relevance between audio samples and captions. In contrast, a recent study used human-assigned relevance ratings, i.e., continuous relevances, for these pairs but did not obtain performance gains in audio-text relevance learning. This work introduces a relevance learning method that utilizes both human-assigned continuous relevance ratings and binary relevances using a combination of a listwise ranking objective and a contrastive learning objective. Experimental results demonstrate the effectiveness of the proposed method, showing improvements in language-based audio retrieval, a downstream task in audio-text relevance learning. In addition, we analyze how properties of the captions or audio clips contribute to the continuous audio-text relevances provided by humans or learned by the machine.
title Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
topic Audio and Speech Processing
url https://arxiv.org/abs/2408.14939