Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moummad, Ilyass, Serizel, Romain, Farrugia, Nicolas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929211867922432
author Moummad, Ilyass
Serizel, Romain
Farrugia, Nicolas
author_facet Moummad, Ilyass
Serizel, Romain
Farrugia, Nicolas
contents Bioacoustic sound event detection allows for better understanding of animal behavior and for better monitoring biodiversity using audio. Deep learning systems can help achieve this goal, however it is difficult to acquire sufficient annotated data to train these systems from scratch. To address this limitation, the Detection and Classification of Acoustic Scenes and Events (DCASE) community has recasted the problem within the framework of few-shot learning and organize an annual challenge for learning to detect animal sounds from only five annotated examples. In this work, we regularize supervised contrastive pre-training to learn features that can transfer well on new target tasks with animal sounds unseen during training, achieving a high F-score of 61.52%(0.48) when no feature adaptation is applied, and an F-score of 68.19%(0.75) when we further adapt the learned features for each new target task. This work aims to lower the entry bar to few-shot bioacoustic sound event detection by proposing a simple and yet effective framework for this task, by also providing open-source code.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08971
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection
Moummad, Ilyass
Serizel, Romain
Farrugia, Nicolas
Sound
Machine Learning
Audio and Speech Processing
Bioacoustic sound event detection allows for better understanding of animal behavior and for better monitoring biodiversity using audio. Deep learning systems can help achieve this goal, however it is difficult to acquire sufficient annotated data to train these systems from scratch. To address this limitation, the Detection and Classification of Acoustic Scenes and Events (DCASE) community has recasted the problem within the framework of few-shot learning and organize an annual challenge for learning to detect animal sounds from only five annotated examples. In this work, we regularize supervised contrastive pre-training to learn features that can transfer well on new target tasks with animal sounds unseen during training, achieving a high F-score of 61.52%(0.48) when no feature adaptation is applied, and an F-score of 68.19%(0.75) when we further adapt the learned features for each new target task. This work aims to lower the entry bar to few-shot bioacoustic sound event detection by proposing a simple and yet effective framework for this task, by also providing open-source code.
title Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2309.08971