CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bartelds, Martijn, Nandi, Ananjan, Doumbouya, Moussa Koulako Bala, Jurafsky, Dan, Hashimoto, Tatsunori, Livescu, Karen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915759240773632
author Bartelds, Martijn
Nandi, Ananjan
Doumbouya, Moussa Koulako Bala
Jurafsky, Dan
Hashimoto, Tatsunori
Livescu, Karen
author_facet Bartelds, Martijn
Nandi, Ananjan
Doumbouya, Moussa Koulako Bala
Jurafsky, Dan
Hashimoto, Tatsunori
Livescu, Karen
contents Modern deep learning models often achieve high overall performance, but consistently fail on specific subgroups. Group distributionally robust optimization (group DRO) addresses this problem by minimizing the worst-group loss, but it fails when group losses misrepresent performance differences between groups. This is common in domains like speech, where the widely used connectionist temporal classification (CTC) loss not only scales with input length but also varies with linguistic and acoustic properties, leading to spurious differences between group losses. We present CTC-DRO, which addresses the shortcomings of the group DRO objective by smoothing the group weight update to prevent overemphasis on consistently high-loss groups, while using input length-matched batching to mitigate CTC's scaling issues. We evaluate CTC-DRO on the task of multilingual automatic speech recognition (ASR) across five language sets from the diverse ML-SUPERB 2.0 benchmark. CTC-DRO consistently outperforms group DRO and CTC-based baseline models, reducing the worst-language error by up to 47.1% and the average error by up to 32.9%. CTC-DRO can be applied to ASR with minimal computational costs, and, while motivated by multilingual ASR, offers the potential for reducing group disparities in other domains with similar challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01777
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
Bartelds, Martijn
Nandi, Ananjan
Doumbouya, Moussa Koulako Bala
Jurafsky, Dan
Hashimoto, Tatsunori
Livescu, Karen
Machine Learning
Computation and Language
Audio and Speech Processing
Modern deep learning models often achieve high overall performance, but consistently fail on specific subgroups. Group distributionally robust optimization (group DRO) addresses this problem by minimizing the worst-group loss, but it fails when group losses misrepresent performance differences between groups. This is common in domains like speech, where the widely used connectionist temporal classification (CTC) loss not only scales with input length but also varies with linguistic and acoustic properties, leading to spurious differences between group losses. We present CTC-DRO, which addresses the shortcomings of the group DRO objective by smoothing the group weight update to prevent overemphasis on consistently high-loss groups, while using input length-matched batching to mitigate CTC's scaling issues. We evaluate CTC-DRO on the task of multilingual automatic speech recognition (ASR) across five language sets from the diverse ML-SUPERB 2.0 benchmark. CTC-DRO consistently outperforms group DRO and CTC-based baseline models, reducing the worst-language error by up to 47.1% and the average error by up to 32.9%. CTC-DRO can be applied to ASR with minimal computational costs, and, while motivated by multilingual ASR, offers the potential for reducing group disparities in other domains with similar challenges.
title CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
topic Machine Learning
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2502.01777