Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kashiwagi, Yosuke, Futami, Hayato, Tsunoo, Emiru, Arora, Siddhant, Watanabe, Shinji
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910493276372992
author Kashiwagi, Yosuke
Futami, Hayato
Tsunoo, Emiru
Arora, Siddhant
Watanabe, Shinji
author_facet Kashiwagi, Yosuke
Futami, Hayato
Tsunoo, Emiru
Arora, Siddhant
Watanabe, Shinji
contents End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language of incoming speech. Since the common scenario is where the language is already known, these models can perform as language-specific by using language information as prompts, which is particularly beneficial for attention-based encoder-decoder architectures. However, the Connectionist Temporal Classification (CTC) approach, which enhances recognition via joint decoding and multi-task training, does not normally incorporate language prompts due to its conditionally independent output tokens. To overcome this, we introduce an encoder prompting technique within the self-conditioned CTC framework, enabling language-specific adaptation of the CTC model in a zero-shot manner. Our method has shown to significantly reduce errors by 28% on average and by 41% on low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
Kashiwagi, Yosuke
Futami, Hayato
Tsunoo, Emiru
Arora, Siddhant
Watanabe, Shinji
Sound
Computation and Language
Audio and Speech Processing
End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language of incoming speech. Since the common scenario is where the language is already known, these models can perform as language-specific by using language information as prompts, which is particularly beneficial for attention-based encoder-decoder architectures. However, the Connectionist Temporal Classification (CTC) approach, which enhances recognition via joint decoding and multi-task training, does not normally incorporate language prompts due to its conditionally independent output tokens. To overcome this, we introduce an encoder prompting technique within the self-conditioned CTC framework, enabling language-specific adaptation of the CTC model in a zero-shot manner. Our method has shown to significantly reduce errors by 28% on average and by 41% on low-resource languages.
title Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2406.12611