The Interspeech 2025 Speech Accessibility Project Challenge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866916869096603648 |
|---|---|
| author | Zheng, Xiuwen Phukon, Bornali Na, Jonghwan Cutrell, Ed Han, Kyu Hasegawa-Johnson, Mark Jiang, Pan-Pan Kuila, Aadhrik Lea, Colin MacDonald, Bob Mantena, Gautam Ravichandran, Venkatesh Sari, Leda Tomanek, Katrin Yoo, Chang D. Zwilling, Chris |
| author_facet | Zheng, Xiuwen Phukon, Bornali Na, Jonghwan Cutrell, Ed Han, Kyu Hasegawa-Johnson, Mark Jiang, Pan-Pan Kuila, Aadhrik Lea, Colin MacDonald, Bob Mantena, Gautam Ravichandran, Venkatesh Sari, Leda Tomanek, Katrin Yoo, Chang D. Zwilling, Chris |
| contents | While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training data. To bridge this gap, the 2025 Interspeech Speech Accessibility Project (SAP) Challenge was launched, utilizing over 400 hours of SAP data collected and transcribed from more than 500 individuals with diverse speech disabilities. Hosted on EvalAI and leveraging the remote evaluation pipeline, the SAP Challenge evaluates submissions based on Word Error Rate and Semantic Score. Consequently, 12 out of 22 valid teams outperformed the whisper-large-v2 baseline in terms of WER, while 17 teams surpassed the baseline on SemScore. Notably, the top team achieved the lowest WER of 8.11\%, and the highest SemScore of 88.44\% at the same time, setting new benchmarks for future ASR systems in recognizing impaired speech. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_22047 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The Interspeech 2025 Speech Accessibility Project Challenge Zheng, Xiuwen Phukon, Bornali Na, Jonghwan Cutrell, Ed Han, Kyu Hasegawa-Johnson, Mark Jiang, Pan-Pan Kuila, Aadhrik Lea, Colin MacDonald, Bob Mantena, Gautam Ravichandran, Venkatesh Sari, Leda Tomanek, Katrin Yoo, Chang D. Zwilling, Chris Artificial Intelligence While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training data. To bridge this gap, the 2025 Interspeech Speech Accessibility Project (SAP) Challenge was launched, utilizing over 400 hours of SAP data collected and transcribed from more than 500 individuals with diverse speech disabilities. Hosted on EvalAI and leveraging the remote evaluation pipeline, the SAP Challenge evaluates submissions based on Word Error Rate and Semantic Score. Consequently, 12 out of 22 valid teams outperformed the whisper-large-v2 baseline in terms of WER, while 17 teams surpassed the baseline on SemScore. Notably, the top team achieved the lowest WER of 8.11\%, and the highest SemScore of 88.44\% at the same time, setting new benchmarks for future ASR systems in recognizing impaired speech. |
| title | The Interspeech 2025 Speech Accessibility Project Challenge |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2507.22047 |