On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910329311592448 |
|---|---|
| author | Heggan, Calum Budgett, Sam Hospedales, Timothy Yaghoobi, Mehrdad |
| author_facet | Heggan, Calum Budgett, Sam Hospedales, Timothy Yaghoobi, Mehrdad |
| contents | In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_01274 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification Heggan, Calum Budgett, Sam Hospedales, Timothy Yaghoobi, Mehrdad Sound Machine Learning Audio and Speech Processing In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks. |
| title | On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2402.01274 |