Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de Groot, Dimme, Patel, Tanvina, Kayande, Devendra, Scharenborg, Odette, Yue, Zhengjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914004059815936
author de Groot, Dimme
Patel, Tanvina
Kayande, Devendra
Scharenborg, Odette
Yue, Zhengjun
author_facet de Groot, Dimme
Patel, Tanvina
Kayande, Devendra
Scharenborg, Odette
Yue, Zhengjun
contents Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17980
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
de Groot, Dimme
Patel, Tanvina
Kayande, Devendra
Scharenborg, Odette
Yue, Zhengjun
Audio and Speech Processing
Sound
Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
title Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2508.17980