Low-resource speech recognition and dialect identification of Irish in a multi-task framework

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lonergan, Liam, Qian, Mengjie, Chiaráin, Neasa Ní, Gobl, Christer, Chasaide, Ailbhe Ní
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913339493318656
author Lonergan, Liam
Qian, Mengjie
Chiaráin, Neasa Ní
Gobl, Christer
Chasaide, Ailbhe Ní
author_facet Lonergan, Liam
Qian, Mengjie
Chiaráin, Neasa Ní
Gobl, Christer
Chasaide, Ailbhe Ní
contents This paper explores the use of Hybrid CTC/Attention encoder-decoder models trained with Intermediate CTC (InterCTC) for Irish (Gaelic) low-resource speech recognition (ASR) and dialect identification (DID). Results are compared to the current best performing models trained for ASR (TDNN-HMM) and DID (ECAPA-TDNN). An optimal InterCTC setting is initially established using a Conformer encoder. This setting is then used to train a model with an E-branchformer encoder and the performance of both architectures are compared. A multi-task fine-tuning approach is adopted for language model (LM) shallow fusion. The experiments yielded an improvement in DID accuracy of 10.8% relative to a baseline ECAPA-TDNN, and WER performance approaching the TDNN-HMM model. This multi-task approach emerges as a promising strategy for Irish low-resource ASR and DID.
format Preprint
id arxiv_https___arxiv_org_abs_2405_01293
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-resource speech recognition and dialect identification of Irish in a multi-task framework
Lonergan, Liam
Qian, Mengjie
Chiaráin, Neasa Ní
Gobl, Christer
Chasaide, Ailbhe Ní
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
This paper explores the use of Hybrid CTC/Attention encoder-decoder models trained with Intermediate CTC (InterCTC) for Irish (Gaelic) low-resource speech recognition (ASR) and dialect identification (DID). Results are compared to the current best performing models trained for ASR (TDNN-HMM) and DID (ECAPA-TDNN). An optimal InterCTC setting is initially established using a Conformer encoder. This setting is then used to train a model with an E-branchformer encoder and the performance of both architectures are compared. A multi-task fine-tuning approach is adopted for language model (LM) shallow fusion. The experiments yielded an improvement in DID accuracy of 10.8% relative to a baseline ECAPA-TDNN, and WER performance approaching the TDNN-HMM model. This multi-task approach emerges as a promising strategy for Irish low-resource ASR and DID.
title Low-resource speech recognition and dialect identification of Irish in a multi-task framework
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.01293