The Interspeech 2025 Speech Accessibility Project Challenge

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zheng, Xiuwen, Phukon, Bornali, Na, Jonghwan, Cutrell, Ed, Han, Kyu, Hasegawa-Johnson, Mark, Jiang, Pan-Pan, Kuila, Aadhrik, Lea, Colin, MacDonald, Bob, Mantena, Gautam, Ravichandran, Venkatesh, Sari, Leda, Tomanek, Katrin, Yoo, Chang D., Zwilling, Chris
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916869096603648
author Zheng, Xiuwen
Phukon, Bornali
Na, Jonghwan
Cutrell, Ed
Han, Kyu
Hasegawa-Johnson, Mark
Jiang, Pan-Pan
Kuila, Aadhrik
Lea, Colin
MacDonald, Bob
Mantena, Gautam
Ravichandran, Venkatesh
Sari, Leda
Tomanek, Katrin
Yoo, Chang D.
Zwilling, Chris
author_facet Zheng, Xiuwen
Phukon, Bornali
Na, Jonghwan
Cutrell, Ed
Han, Kyu
Hasegawa-Johnson, Mark
Jiang, Pan-Pan
Kuila, Aadhrik
Lea, Colin
MacDonald, Bob
Mantena, Gautam
Ravichandran, Venkatesh
Sari, Leda
Tomanek, Katrin
Yoo, Chang D.
Zwilling, Chris
contents While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training data. To bridge this gap, the 2025 Interspeech Speech Accessibility Project (SAP) Challenge was launched, utilizing over 400 hours of SAP data collected and transcribed from more than 500 individuals with diverse speech disabilities. Hosted on EvalAI and leveraging the remote evaluation pipeline, the SAP Challenge evaluates submissions based on Word Error Rate and Semantic Score. Consequently, 12 out of 22 valid teams outperformed the whisper-large-v2 baseline in terms of WER, while 17 teams surpassed the baseline on SemScore. Notably, the top team achieved the lowest WER of 8.11\%, and the highest SemScore of 88.44\% at the same time, setting new benchmarks for future ASR systems in recognizing impaired speech.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22047
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Interspeech 2025 Speech Accessibility Project Challenge
Zheng, Xiuwen
Phukon, Bornali
Na, Jonghwan
Cutrell, Ed
Han, Kyu
Hasegawa-Johnson, Mark
Jiang, Pan-Pan
Kuila, Aadhrik
Lea, Colin
MacDonald, Bob
Mantena, Gautam
Ravichandran, Venkatesh
Sari, Leda
Tomanek, Katrin
Yoo, Chang D.
Zwilling, Chris
Artificial Intelligence
While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training data. To bridge this gap, the 2025 Interspeech Speech Accessibility Project (SAP) Challenge was launched, utilizing over 400 hours of SAP data collected and transcribed from more than 500 individuals with diverse speech disabilities. Hosted on EvalAI and leveraging the remote evaluation pipeline, the SAP Challenge evaluates submissions based on Word Error Rate and Semantic Score. Consequently, 12 out of 22 valid teams outperformed the whisper-large-v2 baseline in terms of WER, while 17 teams surpassed the baseline on SemScore. Notably, the top team achieved the lowest WER of 8.11\%, and the highest SemScore of 88.44\% at the same time, setting new benchmarks for future ASR systems in recognizing impaired speech.
title The Interspeech 2025 Speech Accessibility Project Challenge
topic Artificial Intelligence
url https://arxiv.org/abs/2507.22047