Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiao, Yang, Wang, Siyi, Holden, Eun-Jung, Dang, Ting
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917546488233984
author Xiao, Yang
Wang, Siyi
Holden, Eun-Jung
Dang, Ting
author_facet Xiao, Yang
Wang, Siyi
Holden, Eun-Jung
Dang, Ting
contents Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24863
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
Xiao, Yang
Wang, Siyi
Holden, Eun-Jung
Dang, Ting
Audio and Speech Processing
Sound
Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions.
title Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2605.24863