Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cheng, Yao-Fei, Chen, Li-Wei, Lee, Hung-Shin, Wang, Hsin-Min
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918151064649728
author Cheng, Yao-Fei
Chen, Li-Wei
Lee, Hung-Shin
Wang, Hsin-Min
author_facet Cheng, Yao-Fei
Chen, Li-Wei
Lee, Hung-Shin
Wang, Hsin-Min
contents This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of self-supervised learning (SSL) in low-resource settings, we explore the impact of data volume on the continued pre-training of SSL models. We propose a novel data-selection scheme leveraging a multilingual corpus to augment the limited target language data. This scheme utilizes a language classifier to extract utterance embeddings and employs one-class classifiers to identify utterances phonetically and phonologically proximate to the target languages. Utterances are ranked and selected based on their decision scores, ensuring the inclusion of highly relevant data in the SSL-ASR pipeline. Our experimental results demonstrate the effectiveness of this approach, yielding substantial improvements in ASR performance for both Amis and Seediq. These findings underscore the feasibility and promise of data augmentation through cross-lingual transfer learning for low-resource language ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08872
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
Cheng, Yao-Fei
Chen, Li-Wei
Lee, Hung-Shin
Wang, Hsin-Min
Computation and Language
Sound
Audio and Speech Processing
This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of self-supervised learning (SSL) in low-resource settings, we explore the impact of data volume on the continued pre-training of SSL models. We propose a novel data-selection scheme leveraging a multilingual corpus to augment the limited target language data. This scheme utilizes a language classifier to extract utterance embeddings and employs one-class classifiers to identify utterances phonetically and phonologically proximate to the target languages. Utterances are ranked and selected based on their decision scores, ensuring the inclusion of highly relevant data in the SSL-ASR pipeline. Our experimental results demonstrate the effectiveness of this approach, yielding substantial improvements in ASR performance for both Amis and Seediq. These findings underscore the feasibility and promise of data augmentation through cross-lingual transfer learning for low-resource language ASR.
title Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.08872