Towards Temporally Explainable Dysarthric Speech Clarity Assessment
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913868772540416 |
|---|---|
| author | Park, Seohyun Gupta, Chitralekha Kwan, Michelle Kah Yian Fung, Xinhui Yip, Alexander Wenjun Nanayakkara, Suranga |
| author_facet | Park, Seohyun Gupta, Chitralekha Kwan, Michelle Kah Yian Fung, Xinhui Yip, Alexander Wenjun Nanayakkara, Suranga |
| contents | Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from six speakers reading two passages, annotated by a speech therapist with temporal markers and mispronunciation descriptions. We design a three-stage framework for explainable mispronunciation evaluation: (1) overall clarity scoring, (2) mispronunciation localization, and (3) mispronunciation type classification. We systematically analyze pretrained Automatic Speech Recognition (ASR) models in each stage, assessing their effectiveness in dysarthric speech evaluation (Code available at: https://github.com/augmented-human-lab/interspeech25_speechtherapy, Supplementary webpage: https://apps.ahlab.org/interspeech25_speechtherapy/). Our findings offer clinically relevant insights for automating actionable feedback for pronunciation assessment, which could enable independent practice for patients and help therapists deliver more effective interventions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00454 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Temporally Explainable Dysarthric Speech Clarity Assessment Park, Seohyun Gupta, Chitralekha Kwan, Michelle Kah Yian Fung, Xinhui Yip, Alexander Wenjun Nanayakkara, Suranga Audio and Speech Processing Human-Computer Interaction Sound Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from six speakers reading two passages, annotated by a speech therapist with temporal markers and mispronunciation descriptions. We design a three-stage framework for explainable mispronunciation evaluation: (1) overall clarity scoring, (2) mispronunciation localization, and (3) mispronunciation type classification. We systematically analyze pretrained Automatic Speech Recognition (ASR) models in each stage, assessing their effectiveness in dysarthric speech evaluation (Code available at: https://github.com/augmented-human-lab/interspeech25_speechtherapy, Supplementary webpage: https://apps.ahlab.org/interspeech25_speechtherapy/). Our findings offer clinically relevant insights for automating actionable feedback for pronunciation assessment, which could enable independent practice for patients and help therapists deliver more effective interventions. |
| title | Towards Temporally Explainable Dysarthric Speech Clarity Assessment |
| topic | Audio and Speech Processing Human-Computer Interaction Sound |
| url | https://arxiv.org/abs/2506.00454 |