Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Ling, Zhu, Charlotte, Shi, Shuju
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915548933128192
author Sun, Ling
Zhu, Charlotte
Shi, Shuju
author_facet Sun, Ling
Zhu, Charlotte
Shi, Shuju
contents General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Improve corpus, we show that naive fine-tuning of Whisper reduces average WER but simultaneously widens disparities and disproportionately harms lower-level learners. To address this, we propose two strategies: (i) proficiency-aware multitask learning, jointly optimizing ASR with proficiency classification, and (ii) targeted augmentation, applying spectrogram masking to low-proficiency speech to counter imbalance. These approaches reduce WER by up to 29.4 percent (relative) and insertion/deletion errors by as much as 58.6 percent (relative). Crucially, despite the severe imbalance of the dataset reflecting real-world distributions, both strategies consistently narrow proficiency gaps, advancing equitable ASR for L2 learners.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
Sun, Ling
Zhu, Charlotte
Shi, Shuju
Sound
Artificial Intelligence
68T07 (Primary), 94A12, 68T05 (Secondary)
I.5.4; I.2.7
General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Improve corpus, we show that naive fine-tuning of Whisper reduces average WER but simultaneously widens disparities and disproportionately harms lower-level learners. To address this, we propose two strategies: (i) proficiency-aware multitask learning, jointly optimizing ASR with proficiency classification, and (ii) targeted augmentation, applying spectrogram masking to low-proficiency speech to counter imbalance. These approaches reduce WER by up to 29.4 percent (relative) and insertion/deletion errors by as much as 58.6 percent (relative). Crucially, despite the severe imbalance of the dataset reflecting real-world distributions, both strategies consistently narrow proficiency gaps, advancing equitable ASR for L2 learners.
title Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
topic Sound
Artificial Intelligence
68T07 (Primary), 94A12, 68T05 (Secondary)
I.5.4; I.2.7
url https://arxiv.org/abs/2510.10738