Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915548933128192 |
|---|---|
| author | Sun, Ling Zhu, Charlotte Shi, Shuju |
| author_facet | Sun, Ling Zhu, Charlotte Shi, Shuju |
| contents | General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Improve corpus, we show that naive fine-tuning of Whisper reduces average WER but simultaneously widens disparities and disproportionately harms lower-level learners. To address this, we propose two strategies: (i) proficiency-aware multitask learning, jointly optimizing ASR with proficiency classification, and (ii) targeted augmentation, applying spectrogram masking to low-proficiency speech to counter imbalance. These approaches reduce WER by up to 29.4 percent (relative) and insertion/deletion errors by as much as 58.6 percent (relative). Crucially, despite the severe imbalance of the dataset reflecting real-world distributions, both strategies consistently narrow proficiency gaps, advancing equitable ASR for L2 learners. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_10738 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR Sun, Ling Zhu, Charlotte Shi, Shuju Sound Artificial Intelligence 68T07 (Primary), 94A12, 68T05 (Secondary) I.5.4; I.2.7 General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Improve corpus, we show that naive fine-tuning of Whisper reduces average WER but simultaneously widens disparities and disproportionately harms lower-level learners. To address this, we propose two strategies: (i) proficiency-aware multitask learning, jointly optimizing ASR with proficiency classification, and (ii) targeted augmentation, applying spectrogram masking to low-proficiency speech to counter imbalance. These approaches reduce WER by up to 29.4 percent (relative) and insertion/deletion errors by as much as 58.6 percent (relative). Crucially, despite the severe imbalance of the dataset reflecting real-world distributions, both strategies consistently narrow proficiency gaps, advancing equitable ASR for L2 learners. |
| title | Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR |
| topic | Sound Artificial Intelligence 68T07 (Primary), 94A12, 68T05 (Secondary) I.5.4; I.2.7 |
| url | https://arxiv.org/abs/2510.10738 |