Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gaido, Marco, Papi, Sara, Negri, Matteo, Bentivogli, Luisa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913590712205312
author Gaido, Marco
Papi, Sara
Negri, Matteo
Bentivogli, Luisa
author_facet Gaido, Marco
Papi, Sara
Negri, Matteo
Bentivogli, Luisa
contents The field of natural language processing (NLP) has recently witnessed a transformative shift with the emergence of foundation models, particularly Large Language Models (LLMs) that have revolutionized text-based NLP. This paradigm has extended to other modalities, including speech, where researchers are actively exploring the combination of Speech Foundation Models (SFMs) and LLMs into single, unified models capable of addressing multimodal tasks. Among such tasks, this paper focuses on speech-to-text translation (ST). By examining the published papers on the topic, we propose a unified view of the architectural solutions and training strategies presented so far, highlighting similarities and differences among them. Based on this examination, we not only organize the lessons learned but also show how diverse settings and evaluation approaches hinder the identification of the best-performing solution for each architectural building block and training choice. Lastly, we outline recommendations for future works on the topic aimed at better understanding the strengths and weaknesses of the SFM+LLM solutions for ST.
format Preprint
id arxiv_https___arxiv_org_abs_2402_12025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
Gaido, Marco
Papi, Sara
Negri, Matteo
Bentivogli, Luisa
Computation and Language
The field of natural language processing (NLP) has recently witnessed a transformative shift with the emergence of foundation models, particularly Large Language Models (LLMs) that have revolutionized text-based NLP. This paradigm has extended to other modalities, including speech, where researchers are actively exploring the combination of Speech Foundation Models (SFMs) and LLMs into single, unified models capable of addressing multimodal tasks. Among such tasks, this paper focuses on speech-to-text translation (ST). By examining the published papers on the topic, we propose a unified view of the architectural solutions and training strategies presented so far, highlighting similarities and differences among them. Based on this examination, we not only organize the lessons learned but also show how diverse settings and evaluation approaches hinder the identification of the best-performing solution for each architectural building block and training choice. Lastly, we outline recommendations for future works on the topic aimed at better understanding the strengths and weaknesses of the SFM+LLM solutions for ST.
title Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
topic Computation and Language
url https://arxiv.org/abs/2402.12025