The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, William, Meng, Chutong, Shi, Jiatong, Bartelds, Martijn, Wang, Shih-Heng, Wang, Hsiu-Hsuan, Mosquera, Rafael, Hincapie, Sara, Jurafsky, Dan, Anastasopoulos, Antonis, Lee, Hung-yi, Livescu, Karen, Watanabe, Shinji
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912576916422656
author Chen, William
Meng, Chutong
Shi, Jiatong
Bartelds, Martijn
Wang, Shih-Heng
Wang, Hsiu-Hsuan
Mosquera, Rafael
Hincapie, Sara
Jurafsky, Dan
Anastasopoulos, Antonis
Lee, Hung-yi
Livescu, Karen
Watanabe, Shinji
author_facet Chen, William
Meng, Chutong
Shi, Jiatong
Bartelds, Martijn
Wang, Shih-Heng
Wang, Hsiu-Hsuan
Mosquera, Rafael
Hincapie, Sara
Jurafsky, Dan
Anastasopoulos, Antonis
Lee, Hung-yi
Livescu, Karen
Watanabe, Shinji
contents Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. We construct a new test suite that consists of data from 200+ languages, accents, and dialects to evaluate SOTA multilingual speech models. The challenge also introduces an online evaluation server based on DynaBench, allowing for flexibility in model design and architecture for participants. The challenge received 5 submissions from 3 teams, all of which outperformed our baselines. The best-performing submission achieved an absolute improvement in LID accuracy of 23% and a reduction in CER of 18% when compared to the best baseline on a general multilingual test set. On accented and dialectal data, the best submission obtained 30.2% lower CER and 15.7% higher LID accuracy, showing the importance of community challenges in making speech technologies more inclusive.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07139
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
Chen, William
Meng, Chutong
Shi, Jiatong
Bartelds, Martijn
Wang, Shih-Heng
Wang, Hsiu-Hsuan
Mosquera, Rafael
Hincapie, Sara
Jurafsky, Dan
Anastasopoulos, Antonis
Lee, Hung-yi
Livescu, Karen
Watanabe, Shinji
Computation and Language
Audio and Speech Processing
Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. We construct a new test suite that consists of data from 200+ languages, accents, and dialects to evaluate SOTA multilingual speech models. The challenge also introduces an online evaluation server based on DynaBench, allowing for flexibility in model design and architecture for participants. The challenge received 5 submissions from 3 teams, all of which outperformed our baselines. The best-performing submission achieved an absolute improvement in LID accuracy of 23% and a reduction in CER of 18% when compared to the best baseline on a general multilingual test set. On accented and dialectal data, the best submission obtained 30.2% lower CER and 15.7% higher LID accuracy, showing the importance of community challenges in making speech technologies more inclusive.
title The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2509.07139