Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adila, Aulia, Lestari, Dessi, Purwarianti, Ayu, Tanaya, Dipta, Azizah, Kurniawati, Sakti, Sakriani
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910648126930944
author Adila, Aulia
Lestari, Dessi
Purwarianti, Ayu
Tanaya, Dipta
Azizah, Kurniawati
Sakti, Sakriani
author_facet Adila, Aulia
Lestari, Dessi
Purwarianti, Ayu
Tanaya, Dipta
Azizah, Kurniawati
Sakti, Sakriani
contents An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise conditions (clean and moderate). Building such a model requires a significant amount of training data with diverse speech characteristics. Currently, Indonesian data is dominated by read, formal, and clean speech, leading to a scarcity of Indonesian data with other speech variabilities. To develop Indonesian automatic speech recognition (ASR), we present our research on state-of-the-art speech recognition models, namely Massively Multilingual Speech (MMS) and Whisper, as well as compiling a dataset comprising Indonesian speech with variabilities to facilitate our study. We further investigate the models' predictive ability to transcribe Indonesian speech data across different variability groups. The best results were achieved by the Whisper fine-tuned model across datasets with various characteristics, as indicated by the decrease in word error rate (WER) and character error rate (CER). Moreover, we found that speaking style variability affected model performance the most.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08828
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
Adila, Aulia
Lestari, Dessi
Purwarianti, Ayu
Tanaya, Dipta
Azizah, Kurniawati
Sakti, Sakriani
Computation and Language
Sound
Audio and Speech Processing
An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise conditions (clean and moderate). Building such a model requires a significant amount of training data with diverse speech characteristics. Currently, Indonesian data is dominated by read, formal, and clean speech, leading to a scarcity of Indonesian data with other speech variabilities. To develop Indonesian automatic speech recognition (ASR), we present our research on state-of-the-art speech recognition models, namely Massively Multilingual Speech (MMS) and Whisper, as well as compiling a dataset comprising Indonesian speech with variabilities to facilitate our study. We further investigate the models' predictive ability to transcribe Indonesian speech data across different variability groups. The best results were achieved by the Whisper fine-tuned model across datasets with various characteristics, as indicated by the decrease in word error rate (WER) and character error rate (CER). Moreover, we found that speaking style variability affected model performance the most.
title Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.08828