Scaling few-shot spoken word classification with generative meta-continual learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beyers, Louise, Ziki, Batsirayi Mupamhi, van der Merwe, Ruan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916011497750528
author Beyers, Louise
Ziki, Batsirayi Mupamhi
van der Merwe, Ruan
author_facet Beyers, Louise
Ziki, Batsirayi Mupamhi
van der Merwe, Ruan
contents Few-shot spoken word classification has largely been developed for applications where a small number of classes is considered, and so the potential of larger-scale few-shot spoken word classification remains untapped. This paper investigates the potential of a spoken word classifier to sequentially learn to distinguish between 1000 classes when it is given only five shots per class. We demonstrate that this scaling capability exists by training a model using the Generative Meta-Continual Learning (GeMCL) algorithm and comparing it to repeatedly trained or finetuned baselines. We find that GeMCL produces exceptionally stable performance, and although it does not always outperform a repeatedly fully-finetuned HuBERT model nor a frozen HuBERT model with a repeatedly trained classifier head, it produces comparable performance to the latter while adapting 2000 times faster, having been trained less than half of the data for two orders of magnitude less time.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13075
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling few-shot spoken word classification with generative meta-continual learning
Beyers, Louise
Ziki, Batsirayi Mupamhi
van der Merwe, Ruan
Computation and Language
Artificial Intelligence
Few-shot spoken word classification has largely been developed for applications where a small number of classes is considered, and so the potential of larger-scale few-shot spoken word classification remains untapped. This paper investigates the potential of a spoken word classifier to sequentially learn to distinguish between 1000 classes when it is given only five shots per class. We demonstrate that this scaling capability exists by training a model using the Generative Meta-Continual Learning (GeMCL) algorithm and comparing it to repeatedly trained or finetuned baselines. We find that GeMCL produces exceptionally stable performance, and although it does not always outperform a repeatedly fully-finetuned HuBERT model nor a frozen HuBERT model with a repeatedly trained classifier head, it produces comparable performance to the latter while adapting 2000 times faster, having been trained less than half of the data for two orders of magnitude less time.
title Scaling few-shot spoken word classification with generative meta-continual learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.13075