Self-Supervised Learning for Few-Shot Bird Sound Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moummad, Ilyass, Serizel, Romain, Farrugia, Nicolas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911774050091008
author Moummad, Ilyass
Serizel, Romain
Farrugia, Nicolas
author_facet Moummad, Ilyass
Serizel, Romain
Farrugia, Nicolas
contents Self-supervised learning (SSL) in audio holds significant potential across various domains, particularly in situations where abundant, unlabeled data is readily available at no cost. This is pertinent in bioacoustics, where biologists routinely collect extensive sound datasets from the natural environment. In this study, we demonstrate that SSL is capable of acquiring meaningful representations of bird sounds from audio recordings without the need for annotations. Our experiments showcase that these learned representations exhibit the capacity to generalize to new bird species in few-shot learning (FSL) scenarios. Additionally, we show that selecting windows with high bird activation for self-supervised learning, using a pretrained audio neural network, significantly enhances the quality of the learned representations.
format Preprint
id arxiv_https___arxiv_org_abs_2312_15824
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Self-Supervised Learning for Few-Shot Bird Sound Classification
Moummad, Ilyass
Serizel, Romain
Farrugia, Nicolas
Sound
Machine Learning
Audio and Speech Processing
Self-supervised learning (SSL) in audio holds significant potential across various domains, particularly in situations where abundant, unlabeled data is readily available at no cost. This is pertinent in bioacoustics, where biologists routinely collect extensive sound datasets from the natural environment. In this study, we demonstrate that SSL is capable of acquiring meaningful representations of bird sounds from audio recordings without the need for annotations. Our experiments showcase that these learned representations exhibit the capacity to generalize to new bird species in few-shot learning (FSL) scenarios. Additionally, we show that selecting windows with high bird activation for self-supervised learning, using a pretrained audio neural network, significantly enhances the quality of the learned representations.
title Self-Supervised Learning for Few-Shot Bird Sound Classification
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2312.15824