Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hajal, Karl El, Hermann, Enno, Hovsepyan, Sevada, -Doss, Mathew Magimai.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916772826841088
author Hajal, Karl El
Hermann, Enno
Hovsepyan, Sevada
-Doss, Mathew Magimai.
author_facet Hajal, Karl El
Hermann, Enno
Hovsepyan, Sevada
-Doss, Mathew Magimai.
contents Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our approach extends the Rhythm and Voice (RnV) conversion framework by introducing a syllable-based rhythm modeling method suited for dysarthric speech. We assess its impact on ASR by training LF-MMI models and fine-tuning Whisper on converted speech. Experiments on the Torgo corpus reveal that LF-MMI achieves significant word error rate reductions, especially for more severe cases of dysarthria, while fine-tuning Whisper on converted data has minimal effect on its performance. These results highlight the potential of unsupervised rhythm and voice conversion for dysarthric ASR. Code available at: https://github.com/idiap/RnV
format Preprint
id arxiv_https___arxiv_org_abs_2506_01618
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
Hajal, Karl El
Hermann, Enno
Hovsepyan, Sevada
-Doss, Mathew Magimai.
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our approach extends the Rhythm and Voice (RnV) conversion framework by introducing a syllable-based rhythm modeling method suited for dysarthric speech. We assess its impact on ASR by training LF-MMI models and fine-tuning Whisper on converted speech. Experiments on the Torgo corpus reveal that LF-MMI achieves significant word error rate reductions, especially for more severe cases of dysarthria, while fine-tuning Whisper on converted data has minimal effect on its performance. These results highlight the potential of unsupervised rhythm and voice conversion for dysarthric ASR. Code available at: https://github.com/idiap/RnV
title Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2506.01618