Zero-resource Speech Translation and Recognition with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mundnich, Karel, Niu, Xing, Mathur, Prashant, Ronanki, Srikanth, Houston, Brady, Elluru, Veera Raghavendra, Das, Nilaksh, Hou, Zejiang, Huybrechts, Goeric, Bhatia, Anshu, Garcia-Romero, Daniel, Han, Kyu J., Kirchhoff, Katrin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910768437395456
author Mundnich, Karel
Niu, Xing
Mathur, Prashant
Ronanki, Srikanth
Houston, Brady
Elluru, Veera Raghavendra
Das, Nilaksh
Hou, Zejiang
Huybrechts, Goeric
Bhatia, Anshu
Garcia-Romero, Daniel
Han, Kyu J.
Kirchhoff, Katrin
author_facet Mundnich, Karel
Niu, Xing
Mathur, Prashant
Ronanki, Srikanth
Houston, Brady
Elluru, Veera Raghavendra
Das, Nilaksh
Hou, Zejiang
Huybrechts, Goeric
Bhatia, Anshu
Garcia-Romero, Daniel
Han, Kyu J.
Kirchhoff, Katrin
contents Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2\%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18566
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-resource Speech Translation and Recognition with LLMs
Mundnich, Karel
Niu, Xing
Mathur, Prashant
Ronanki, Srikanth
Houston, Brady
Elluru, Veera Raghavendra
Das, Nilaksh
Hou, Zejiang
Huybrechts, Goeric
Bhatia, Anshu
Garcia-Romero, Daniel
Han, Kyu J.
Kirchhoff, Katrin
Computation and Language
Audio and Speech Processing
Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2\%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language.
title Zero-resource Speech Translation and Recognition with LLMs
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2412.18566