Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Primus, Paul, Widmer, Gerhard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914855516110848
author Primus, Paul
Widmer, Gerhard
author_facet Primus, Paul
Widmer, Gerhard
contents Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that utilizes audio metadata as an additional clue to understand the content of audio signals before matching them with textual queries. We experimented with metadata often attached to audio recordings, such as keywords and natural-language descriptions, and we investigated late and mid-level fusion strategies to merge audio and metadata. Our hybrid approach with keyword metadata and late fusion improved the retrieval performance over a content-based baseline by 2.36 and 3.69 pp. mAP@10 on the ClothoV2 and AudioCaps benchmarks, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15897
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
Primus, Paul
Widmer, Gerhard
Audio and Speech Processing
Machine Learning
Sound
Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that utilizes audio metadata as an additional clue to understand the content of audio signals before matching them with textual queries. We experimented with metadata often attached to audio recordings, such as keywords and natural-language descriptions, and we investigated late and mid-level fusion strategies to merge audio and metadata. Our hybrid approach with keyword metadata and late fusion improved the retrieval performance over a content-based baseline by 2.36 and 3.69 pp. mAP@10 on the ClothoV2 and AudioCaps benchmarks, respectively.
title Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2406.15897