Configurable Multilingual ASR with Speech Summary Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Harrison, Fung, Ivan, Zhu, Yingke, Samarakoon, Lahiru
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910635383586816
author Zhu, Harrison
Fung, Ivan
Zhu, Yingke
Samarakoon, Lahiru
author_facet Zhu, Harrison
Fung, Ivan
Zhu, Yingke
Samarakoon, Lahiru
contents Approximately half of the world's population is multilingual, making multilingual ASR (MASR) essential. Deploying multiple monolingual models is challenging when the ground-truth language is unknown in advance. This motivates research efforts on configurable multilingual MASR models that can be prompted manually or adapted automatically to recognise specific languages. In this paper, we present the Configurable MASR model with Summary Vector (csvMASR), a novel architecture designed to enhance configurability. Our approach leverages adapters and introduces speech summary vector representations, inspired by conversational summary representations in speech diarization, to combine outputs from language-specific components at the utterance level. We also incorporate an auxiliary language classification loss to enhance configurability. Using data from 7 languages in the Multilingual Librispeech (MLS) dataset, csvMASR outperforms existing MASR models and reduces the word error rate (WER) from 10.33\% to 9.95\% when compared with the baseline. Additionally, csvMASR demonstrates superior performance in language classification and prompting tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04478
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Configurable Multilingual ASR with Speech Summary Representations
Zhu, Harrison
Fung, Ivan
Zhu, Yingke
Samarakoon, Lahiru
Sound
Computation and Language
Audio and Speech Processing
Approximately half of the world's population is multilingual, making multilingual ASR (MASR) essential. Deploying multiple monolingual models is challenging when the ground-truth language is unknown in advance. This motivates research efforts on configurable multilingual MASR models that can be prompted manually or adapted automatically to recognise specific languages. In this paper, we present the Configurable MASR model with Summary Vector (csvMASR), a novel architecture designed to enhance configurability. Our approach leverages adapters and introduces speech summary vector representations, inspired by conversational summary representations in speech diarization, to combine outputs from language-specific components at the utterance level. We also incorporate an auxiliary language classification loss to enhance configurability. Using data from 7 languages in the Multilingual Librispeech (MLS) dataset, csvMASR outperforms existing MASR models and reduces the word error rate (WER) from 10.33\% to 9.95\% when compared with the baseline. Additionally, csvMASR demonstrates superior performance in language classification and prompting tasks.
title Configurable Multilingual ASR with Speech Summary Representations
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2410.04478