Stuttering-Aware Automatic Speech Recognition for Indonesian Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muhammad, Fadhil, Djuliansah, Alwin, Hamzah, Adrian Aryaputra, Azizah, Kurniawati
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912822460416000
author Muhammad, Fadhil
Djuliansah, Alwin
Hamzah, Adrian Aryaputra
Azizah, Kurniawati
author_facet Muhammad, Fadhil
Djuliansah, Alwin
Hamzah, Adrian Aryaputra
Azizah, Kurniawati
contents Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like Indonesian where specialized datasets are virtually non-existent. To overcome this scarcity, we propose a data augmentation framework that generates synthetic stuttered audio by injecting repetitions and prolongations into fluent text through a combination of rule-based transformations and large language models followed by text-to-speech synthesis. We apply this synthetic data to fine-tune a pre-trained Indonesian Whisper model using transfer learning, enabling the architecture to adapt to dysfluent acoustic patterns without requiring large-scale real-world recordings. Our experiments demonstrate that this targeted synthetic exposure consistently reduces recognition errors on stuttered speech while maintaining performance on fluent segments, validating the utility of synthetic data pipelines for developing more inclusive speech technologies in under-represented languages.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03727
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stuttering-Aware Automatic Speech Recognition for Indonesian Language
Muhammad, Fadhil
Djuliansah, Alwin
Hamzah, Adrian Aryaputra
Azizah, Kurniawati
Computation and Language
Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like Indonesian where specialized datasets are virtually non-existent. To overcome this scarcity, we propose a data augmentation framework that generates synthetic stuttered audio by injecting repetitions and prolongations into fluent text through a combination of rule-based transformations and large language models followed by text-to-speech synthesis. We apply this synthetic data to fine-tune a pre-trained Indonesian Whisper model using transfer learning, enabling the architecture to adapt to dysfluent acoustic patterns without requiring large-scale real-world recordings. Our experiments demonstrate that this targeted synthetic exposure consistently reduces recognition errors on stuttered speech while maintaining performance on fluent segments, validating the utility of synthetic data pipelines for developing more inclusive speech technologies in under-represented languages.
title Stuttering-Aware Automatic Speech Recognition for Indonesian Language
topic Computation and Language
url https://arxiv.org/abs/2601.03727