Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Das, Shoutrik, Singh, Nishant, Gangwar, Arjun, Umesh, S
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909653058715648
author Das, Shoutrik
Singh, Nishant
Gangwar, Arjun
Umesh, S
author_facet Das, Shoutrik
Singh, Nishant
Gangwar, Arjun
Umesh, S
contents Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech conversion techniques. In this work, we investigate the utility and limitations of self-supervised learning (SSL) features and their quantized representations as an alternative to mel-spectrograms for speech generation. Additionally, we explore methods to mitigate speaker variability by generating clean speech in a single-speaker voice using features extracted from WavLM. To this end, we propose a fully non-autoregressive approach that leverages Conditional Flow Matching (CFM) with Diffusion Transformers to learn a direct mapping from dysarthric to clean speech. Our findings highlight the effectiveness of discrete acoustic units in improving intelligibility while achieving faster convergence compared to traditional mel-spectrogram-based approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
Das, Shoutrik
Singh, Nishant
Gangwar, Arjun
Umesh, S
Sound
Artificial Intelligence
Audio and Speech Processing
Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech conversion techniques. In this work, we investigate the utility and limitations of self-supervised learning (SSL) features and their quantized representations as an alternative to mel-spectrograms for speech generation. Additionally, we explore methods to mitigate speaker variability by generating clean speech in a single-speaker voice using features extracted from WavLM. To this end, we propose a fully non-autoregressive approach that leverages Conditional Flow Matching (CFM) with Diffusion Transformers to learn a direct mapping from dysarthric to clean speech. Our findings highlight the effectiveness of discrete acoustic units in improving intelligibility while achieving faster convergence compared to traditional mel-spectrogram-based approaches.
title Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.16127