Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chaparala, Kaavya, Zarrella, Guido, Fischer, Bruce Torres, Kimura, Larry, Jones, Oiwi Parker
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909160421982208
author Chaparala, Kaavya
Zarrella, Guido
Fischer, Bruce Torres
Kimura, Larry
Jones, Oiwi Parker
author_facet Chaparala, Kaavya
Zarrella, Guido
Fischer, Bruce Torres
Kimura, Larry
Jones, Oiwi Parker
contents In this paper we address the challenge of improving Automatic Speech Recognition (ASR) for a low-resource language, Hawaiian, by incorporating large amounts of independent text data into an ASR foundation model, Whisper. To do this, we train an external language model (LM) on ~1.5M words of Hawaiian text. We then use the LM to rescore Whisper and compute word error rates (WERs) on a manually curated test set of labeled Hawaiian data. As a baseline, we use Whisper without an external LM. Experimental results reveal a small but significant improvement in WER when ASR outputs are rescored with a Hawaiian LM. The results support leveraging all available data in the development of ASR systems for underrepresented languages.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03073
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian
Chaparala, Kaavya
Zarrella, Guido
Fischer, Bruce Torres
Kimura, Larry
Jones, Oiwi Parker
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
In this paper we address the challenge of improving Automatic Speech Recognition (ASR) for a low-resource language, Hawaiian, by incorporating large amounts of independent text data into an ASR foundation model, Whisper. To do this, we train an external language model (LM) on ~1.5M words of Hawaiian text. We then use the LM to rescore Whisper and compute word error rates (WERs) on a manually curated test set of labeled Hawaiian data. As a baseline, we use Whisper without an external LM. Experimental results reveal a small but significant improvement in WER when ASR outputs are rescored with a Hawaiian LM. The results support leveraging all available data in the development of ASR systems for underrepresented languages.
title Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2404.03073