Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917546488233984 |
|---|---|
| author | Xiao, Yang Wang, Siyi Holden, Eun-Jung Dang, Ting |
| author_facet | Xiao, Yang Wang, Siyi Holden, Eun-Jung Dang, Ting |
| contents | Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_24863 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems Xiao, Yang Wang, Siyi Holden, Eun-Jung Dang, Ting Audio and Speech Processing Sound Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions. |
| title | Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2605.24863 |