Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Greif, Jonathan, Schmid, Florian, Primus, Paul, Widmer, Gerhard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911998016487424
author Greif, Jonathan
Schmid, Florian
Primus, Paul
Widmer, Gerhard
author_facet Greif, Jonathan
Schmid, Florian
Primus, Paul
Widmer, Gerhard
contents Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sound concepts through voice, QBV offers the more intuitive and convenient approach compared to text-based search. To fully leverage QBV, developing robust audio feature representations for both the vocal imitation and the original sound is crucial. In this paper, we present a new system for QBV that utilizes the feature extraction capabilities of Convolutional Neural Networks pre-trained with large-scale general-purpose audio datasets. We integrate these pre-trained models into a dual encoder architecture and fine-tune them end-to-end using contrastive learning. A distinctive aspect of our proposed method is the fine-tuning strategy of pre-trained models using an adapted NT-Xent loss for contrastive learning, creating a shared embedding space for reference recordings and vocal imitations. The proposed system significantly enhances audio retrieval performance, establishing a new state of the art on both coarse- and fine-grained QBV tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11638
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
Greif, Jonathan
Schmid, Florian
Primus, Paul
Widmer, Gerhard
Audio and Speech Processing
Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sound concepts through voice, QBV offers the more intuitive and convenient approach compared to text-based search. To fully leverage QBV, developing robust audio feature representations for both the vocal imitation and the original sound is crucial. In this paper, we present a new system for QBV that utilizes the feature extraction capabilities of Convolutional Neural Networks pre-trained with large-scale general-purpose audio datasets. We integrate these pre-trained models into a dual encoder architecture and fine-tune them end-to-end using contrastive learning. A distinctive aspect of our proposed method is the fine-tuning strategy of pre-trained models using an adapted NT-Xent loss for contrastive learning, creating a shared embedding space for reference recordings and vocal imitations. The proposed system significantly enhances audio retrieval performance, establishing a new state of the art on both coarse- and fine-grained QBV tasks.
title Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
topic Audio and Speech Processing
url https://arxiv.org/abs/2408.11638