On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Heggan, Calum, Budgett, Sam, Hospedales, Timothy, Yaghoobi, Mehrdad
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910329311592448
author Heggan, Calum
Budgett, Sam
Hospedales, Timothy
Yaghoobi, Mehrdad
author_facet Heggan, Calum
Budgett, Sam
Hospedales, Timothy
Yaghoobi, Mehrdad
contents In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01274
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
Heggan, Calum
Budgett, Sam
Hospedales, Timothy
Yaghoobi, Mehrdad
Sound
Machine Learning
Audio and Speech Processing
In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks.
title On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2402.01274