Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mujtaba, Dena, Mahapatra, Nihar
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912545964556288
author Mujtaba, Dena
Mahapatra, Nihar
author_facet Mujtaba, Dena
Mahapatra, Nihar
contents Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven technologies inaccessible to people who stutter. The variability of disfluencies across speakers and contexts further complicates ASR training, compounded by limited annotated stuttered speech data. In this paper, we investigate fine-tuning ASRs for stuttered speech, comparing generalized models (trained across multiple speakers) to personalized models tailored to individual speech characteristics. Using a diverse range of voice-AI scenarios, including virtual assistants and video interviews, we evaluate how personalization affects transcription accuracy. Our findings show that personalized ASRs significantly reduce word error rates, especially in spontaneous speech, highlighting the potential of tailored models for more inclusive voice technologies.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
Mujtaba, Dena
Mahapatra, Nihar
Sound
Audio and Speech Processing
Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven technologies inaccessible to people who stutter. The variability of disfluencies across speakers and contexts further complicates ASR training, compounded by limited annotated stuttered speech data. In this paper, we investigate fine-tuning ASRs for stuttered speech, comparing generalized models (trained across multiple speakers) to personalized models tailored to individual speech characteristics. Using a diverse range of voice-AI scenarios, including virtual assistants and video interviews, we evaluate how personalization affects transcription accuracy. Our findings show that personalized ASRs significantly reduce word error rates, especially in spontaneous speech, highlighting the potential of tailored models for more inclusive voice technologies.
title Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00853