Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Papi, Sara, Gilabert, Javier Garcia, Hopton, Zachary, Zouhar, Vilém, Escolano, Carlos, Gállego, Gerard I., Iranzo-Sánchez, Jorge, Kim, Ahrii, Macháček, Dominik, Schmidtova, Patricia, Züfle, Maike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917435786919936
author Papi, Sara
Gilabert, Javier Garcia
Hopton, Zachary
Zouhar, Vilém
Escolano, Carlos
Gállego, Gerard I.
Iranzo-Sánchez, Jorge
Kim, Ahrii
Macháček, Dominik
Schmidtova, Patricia
Züfle, Maike
author_facet Papi, Sara
Gilabert, Javier Garcia
Hopton, Zachary
Zouhar, Vilém
Escolano, Carlos
Gállego, Gerard I.
Iranzo-Sánchez, Jorge
Kim, Ahrii
Macháček, Dominik
Schmidtova, Patricia
Züfle, Maike
contents As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-to-text translation (ST) and other downstream tasks, bypassing traditional transcription-based pipelines. Whether this integration improves ST quality over established cascaded architectures, however, remains an open question. We present Hearing to Translate, the first comprehensive test suite rigorously benchmarking 6 state-of-the-art SpeechLLMs against 16 strong direct and cascade systems that couple leading speech foundation models (SFM), with multilingual LLMs. Our analysis spans 16 benchmarks, 13 language pairs, and 9 challenging conditions, including disfluent, noisy, and long-form speech. Across this extensive evaluation, we find that cascaded systems remain the most reliable solution overall, but most recent SpeechLLMs can match or even outperform cascades in various settings while SFMs lag behind both, highlighting that integrating an LLM, either within the model or in a pipeline, is essential for high-quality speech translation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16378
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
Papi, Sara
Gilabert, Javier Garcia
Hopton, Zachary
Zouhar, Vilém
Escolano, Carlos
Gállego, Gerard I.
Iranzo-Sánchez, Jorge
Kim, Ahrii
Macháček, Dominik
Schmidtova, Patricia
Züfle, Maike
Computation and Language
Artificial Intelligence
Sound
As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-to-text translation (ST) and other downstream tasks, bypassing traditional transcription-based pipelines. Whether this integration improves ST quality over established cascaded architectures, however, remains an open question. We present Hearing to Translate, the first comprehensive test suite rigorously benchmarking 6 state-of-the-art SpeechLLMs against 16 strong direct and cascade systems that couple leading speech foundation models (SFM), with multilingual LLMs. Our analysis spans 16 benchmarks, 13 language pairs, and 9 challenging conditions, including disfluent, noisy, and long-form speech. Across this extensive evaluation, we find that cascaded systems remain the most reliable solution overall, but most recent SpeechLLMs can match or even outperform cascades in various settings while SFMs lag behind both, highlighting that integrating an LLM, either within the model or in a pipeline, is essential for high-quality speech translation.
title Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
topic Computation and Language
Artificial Intelligence
Sound
url https://arxiv.org/abs/2512.16378