Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Shao-Syuan, Huang, Kuan-Po, Liu, Andy T., Lee, Hung-yi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917876059865088
author Huang, Shao-Syuan
Huang, Kuan-Po
Liu, Andy T.
Lee, Hung-yi
author_facet Huang, Shao-Syuan
Huang, Kuan-Po
Liu, Andy T.
Lee, Hung-yi
contents Multilingual Automatic Speech Recognition (ASR) aims to recognize and transcribe speech from multiple languages within a single system. Whisper, one of the most advanced ASR models, excels in this domain by handling 99 languages effectively, leveraging a vast amount of data and incorporating language tags as prefixes to guide the recognition process. However, despite its success, Whisper struggles with unseen languages, those not included in its pre-training. Motivated by the observation that many languages share linguistic characteristics, we propose methods that exploit these relationships to enhance ASR performance on unseen languages. Specifically, we introduce a weighted sum method, which computes a weighted sum of the embeddings of language tags, using Whisper's predicted language probabilities. In addition, we develop a predictor-based approach that refines the weighted sum embedding to more closely approximate the true embedding for unseen languages. Experimental results demonstrate substantial improvements in ASR performance, both in zero-shot and fine-tuning settings. Our proposed methods outperform baseline approaches, providing an effective solution for addressing unseen languages in multilingual ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16474
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
Huang, Shao-Syuan
Huang, Kuan-Po
Liu, Andy T.
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Multilingual Automatic Speech Recognition (ASR) aims to recognize and transcribe speech from multiple languages within a single system. Whisper, one of the most advanced ASR models, excels in this domain by handling 99 languages effectively, leveraging a vast amount of data and incorporating language tags as prefixes to guide the recognition process. However, despite its success, Whisper struggles with unseen languages, those not included in its pre-training. Motivated by the observation that many languages share linguistic characteristics, we propose methods that exploit these relationships to enhance ASR performance on unseen languages. Specifically, we introduce a weighted sum method, which computes a weighted sum of the embeddings of language tags, using Whisper's predicted language probabilities. In addition, we develop a predictor-based approach that refines the weighted sum embedding to more closely approximate the true embedding for unseen languages. Experimental results demonstrate substantial improvements in ASR performance, both in zero-shot and fine-tuning settings. Our proposed methods outperform baseline approaches, providing an effective solution for addressing unseen languages in multilingual ASR.
title Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2412.16474