Improved Dysarthric Speech to Text Conversion via TTS Personalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mihajlik, Péter, Székely, Éva, Barta, Piroska, Kádár, Máté Soma, Dobsinszki, Gergely, Tóth, László
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908482506063872
author Mihajlik, Péter
Székely, Éva
Barta, Piroska
Kádár, Máté Soma
Dobsinszki, Gergely
Tóth, László
author_facet Mihajlik, Péter
Székely, Éva
Barta, Piroska
Kádár, Máté Soma
Dobsinszki, Gergely
Tóth, László
contents We present a case study on developing a customized speech-to-text system for a Hungarian speaker with severe dysarthria. State-of-the-art automatic speech recognition (ASR) models struggle with zero-shot transcription of dysarthric speech, yielding high error rates. To improve performance with limited real dysarthric data, we fine-tune an ASR model using synthetic speech generated via a personalized text-to-speech (TTS) system. We introduce a method for generating synthetic dysarthric speech with controlled severity by leveraging premorbidity recordings of the given speaker and speaker embedding interpolation, enabling ASR fine-tuning on a continuum of impairments. Fine-tuning on both real and synthetic dysarthric speech reduces the character error rate (CER) from 36-51% (zero-shot) to 7.3%. Our monolingual FastConformer_Hu ASR model significantly outperforms Whisper-turbo when fine-tuned on the same data, and the inclusion of synthetic speech contributes to an 18% relative CER reduction. These results highlight the potential of personalized ASR systems for improving accessibility for individuals with severe speech impairments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improved Dysarthric Speech to Text Conversion via TTS Personalization
Mihajlik, Péter
Székely, Éva
Barta, Piroska
Kádár, Máté Soma
Dobsinszki, Gergely
Tóth, László
Sound
Human-Computer Interaction
We present a case study on developing a customized speech-to-text system for a Hungarian speaker with severe dysarthria. State-of-the-art automatic speech recognition (ASR) models struggle with zero-shot transcription of dysarthric speech, yielding high error rates. To improve performance with limited real dysarthric data, we fine-tune an ASR model using synthetic speech generated via a personalized text-to-speech (TTS) system. We introduce a method for generating synthetic dysarthric speech with controlled severity by leveraging premorbidity recordings of the given speaker and speaker embedding interpolation, enabling ASR fine-tuning on a continuum of impairments. Fine-tuning on both real and synthetic dysarthric speech reduces the character error rate (CER) from 36-51% (zero-shot) to 7.3%. Our monolingual FastConformer_Hu ASR model significantly outperforms Whisper-turbo when fine-tuned on the same data, and the inclusion of synthetic speech contributes to an 18% relative CER reduction. These results highlight the potential of personalized ASR systems for improving accessibility for individuals with severe speech impairments.
title Improved Dysarthric Speech to Text Conversion via TTS Personalization
topic Sound
Human-Computer Interaction
url https://arxiv.org/abs/2508.06391