Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dong, Lukuan, Qin, Donghong, Bai, Fengbo, Song, Fanhua, Liu, Yan, Xu, Chen, Ou, Zhijian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914949806161920
author Dong, Lukuan
Qin, Donghong
Bai, Fengbo
Song, Fanhua
Liu, Yan
Xu, Chen
Ou, Zhijian
author_facet Dong, Lukuan
Qin, Donghong
Bai, Fengbo
Song, Fanhua
Liu, Yan
Xu, Chen
Ou, Zhijian
contents The mainstream automatic speech recognition (ASR) technology usually requires hundreds to thousands of hours of annotated speech data. Three approaches to low-resourced ASR are phoneme or subword based supervised pre-training, and self-supervised pre-training over multilingual data. The Iu Mien language is the main ethnic language of the Yao ethnic group in China and is low-resourced in the sense that the annotated speech is very limited. With less than 10 hours of transcribed Iu Mien language, this paper investigates and compares the three approaches for Iu Mien speech recognition. Our experiments are based on the recently released, three backbone models pretrained over the 10 languages from the CommonVoice dataset (CV-Lang10), which correspond to the three approaches for low-resourced ASR. It is found that phoneme supervision can achieve better results compared to subword supervision and self-supervision, thereby providing higher data-efficiency. Particularly, the Whistle models, i.e., obtained by the weakly-supervised phoneme-based multilingual pre-training, obtain the most competitive results.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
Dong, Lukuan
Qin, Donghong
Bai, Fengbo
Song, Fanhua
Liu, Yan
Xu, Chen
Ou, Zhijian
Sound
Computation and Language
Audio and Speech Processing
The mainstream automatic speech recognition (ASR) technology usually requires hundreds to thousands of hours of annotated speech data. Three approaches to low-resourced ASR are phoneme or subword based supervised pre-training, and self-supervised pre-training over multilingual data. The Iu Mien language is the main ethnic language of the Yao ethnic group in China and is low-resourced in the sense that the annotated speech is very limited. With less than 10 hours of transcribed Iu Mien language, this paper investigates and compares the three approaches for Iu Mien speech recognition. Our experiments are based on the recently released, three backbone models pretrained over the 10 languages from the CommonVoice dataset (CV-Lang10), which correspond to the three approaches for low-resourced ASR. It is found that phoneme supervision can achieve better results compared to subword supervision and self-supervision, thereby providing higher data-efficiency. Particularly, the Whistle models, i.e., obtained by the weakly-supervised phoneme-based multilingual pre-training, obtain the most competitive results.
title Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2407.13292