Towards Temporally Explainable Dysarthric Speech Clarity Assessment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Seohyun, Gupta, Chitralekha, Kwan, Michelle Kah Yian, Fung, Xinhui, Yip, Alexander Wenjun, Nanayakkara, Suranga
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913868772540416
author Park, Seohyun
Gupta, Chitralekha
Kwan, Michelle Kah Yian
Fung, Xinhui
Yip, Alexander Wenjun
Nanayakkara, Suranga
author_facet Park, Seohyun
Gupta, Chitralekha
Kwan, Michelle Kah Yian
Fung, Xinhui
Yip, Alexander Wenjun
Nanayakkara, Suranga
contents Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from six speakers reading two passages, annotated by a speech therapist with temporal markers and mispronunciation descriptions. We design a three-stage framework for explainable mispronunciation evaluation: (1) overall clarity scoring, (2) mispronunciation localization, and (3) mispronunciation type classification. We systematically analyze pretrained Automatic Speech Recognition (ASR) models in each stage, assessing their effectiveness in dysarthric speech evaluation (Code available at: https://github.com/augmented-human-lab/interspeech25_speechtherapy, Supplementary webpage: https://apps.ahlab.org/interspeech25_speechtherapy/). Our findings offer clinically relevant insights for automating actionable feedback for pronunciation assessment, which could enable independent practice for patients and help therapists deliver more effective interventions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Temporally Explainable Dysarthric Speech Clarity Assessment
Park, Seohyun
Gupta, Chitralekha
Kwan, Michelle Kah Yian
Fung, Xinhui
Yip, Alexander Wenjun
Nanayakkara, Suranga
Audio and Speech Processing
Human-Computer Interaction
Sound
Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from six speakers reading two passages, annotated by a speech therapist with temporal markers and mispronunciation descriptions. We design a three-stage framework for explainable mispronunciation evaluation: (1) overall clarity scoring, (2) mispronunciation localization, and (3) mispronunciation type classification. We systematically analyze pretrained Automatic Speech Recognition (ASR) models in each stage, assessing their effectiveness in dysarthric speech evaluation (Code available at: https://github.com/augmented-human-lab/interspeech25_speechtherapy, Supplementary webpage: https://apps.ahlab.org/interspeech25_speechtherapy/). Our findings offer clinically relevant insights for automating actionable feedback for pronunciation assessment, which could enable independent practice for patients and help therapists deliver more effective interventions.
title Towards Temporally Explainable Dysarthric Speech Clarity Assessment
topic Audio and Speech Processing
Human-Computer Interaction
Sound
url https://arxiv.org/abs/2506.00454