CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915759240773632 |
|---|---|
| author | Bartelds, Martijn Nandi, Ananjan Doumbouya, Moussa Koulako Bala Jurafsky, Dan Hashimoto, Tatsunori Livescu, Karen |
| author_facet | Bartelds, Martijn Nandi, Ananjan Doumbouya, Moussa Koulako Bala Jurafsky, Dan Hashimoto, Tatsunori Livescu, Karen |
| contents | Modern deep learning models often achieve high overall performance, but consistently fail on specific subgroups. Group distributionally robust optimization (group DRO) addresses this problem by minimizing the worst-group loss, but it fails when group losses misrepresent performance differences between groups. This is common in domains like speech, where the widely used connectionist temporal classification (CTC) loss not only scales with input length but also varies with linguistic and acoustic properties, leading to spurious differences between group losses. We present CTC-DRO, which addresses the shortcomings of the group DRO objective by smoothing the group weight update to prevent overemphasis on consistently high-loss groups, while using input length-matched batching to mitigate CTC's scaling issues. We evaluate CTC-DRO on the task of multilingual automatic speech recognition (ASR) across five language sets from the diverse ML-SUPERB 2.0 benchmark. CTC-DRO consistently outperforms group DRO and CTC-based baseline models, reducing the worst-language error by up to 47.1% and the average error by up to 32.9%. CTC-DRO can be applied to ASR with minimal computational costs, and, while motivated by multilingual ASR, offers the potential for reducing group disparities in other domains with similar challenges. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_01777 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition Bartelds, Martijn Nandi, Ananjan Doumbouya, Moussa Koulako Bala Jurafsky, Dan Hashimoto, Tatsunori Livescu, Karen Machine Learning Computation and Language Audio and Speech Processing Modern deep learning models often achieve high overall performance, but consistently fail on specific subgroups. Group distributionally robust optimization (group DRO) addresses this problem by minimizing the worst-group loss, but it fails when group losses misrepresent performance differences between groups. This is common in domains like speech, where the widely used connectionist temporal classification (CTC) loss not only scales with input length but also varies with linguistic and acoustic properties, leading to spurious differences between group losses. We present CTC-DRO, which addresses the shortcomings of the group DRO objective by smoothing the group weight update to prevent overemphasis on consistently high-loss groups, while using input length-matched batching to mitigate CTC's scaling issues. We evaluate CTC-DRO on the task of multilingual automatic speech recognition (ASR) across five language sets from the diverse ML-SUPERB 2.0 benchmark. CTC-DRO consistently outperforms group DRO and CTC-based baseline models, reducing the worst-language error by up to 47.1% and the average error by up to 32.9%. CTC-DRO can be applied to ASR with minimal computational costs, and, while motivated by multilingual ASR, offers the potential for reducing group disparities in other domains with similar challenges. |
| title | CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition |
| topic | Machine Learning Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2502.01777 |