Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pareras, Oriol, Gállego, Gerard I., Costa, Federico, España-Bonet, Cristina, Hernando, Javier
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918473569927168
author Pareras, Oriol
Gállego, Gerard I.
Costa, Federico
España-Bonet, Cristina
Hernando, Javier
author_facet Pareras, Oriol
Gállego, Gerard I.
Costa, Federico
España-Bonet, Cristina
Hernando, Javier
contents Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then translate it. CoT typically outperforms direct prompting primarily because it can exploit abundant Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) datasets to explicitly model its steps. In this paper, we systematically compare CoT and Direct prompting under increasing amounts of S2TT data. To this end, we pseudo-label an ASR corpus by translating its transcriptions into six European languages, and train LLM-based S2TT systems with both prompting strategies at different data scales. Our results show that Direct improves more consistently as the amount of data increases, suggesting that it may become a more effective approach as larger S2TT resources are created.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
Pareras, Oriol
Gállego, Gerard I.
Costa, Federico
España-Bonet, Cristina
Hernando, Javier
Computation and Language
Sound
Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then translate it. CoT typically outperforms direct prompting primarily because it can exploit abundant Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) datasets to explicitly model its steps. In this paper, we systematically compare CoT and Direct prompting under increasing amounts of S2TT data. To this end, we pseudo-label an ASR corpus by translating its transcriptions into six European languages, and train LLM-based S2TT systems with both prompting strategies at different data scales. Our results show that Direct improves more consistently as the amount of data increases, suggesting that it may become a more effective approach as larger S2TT resources are created.
title Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
topic Computation and Language
Sound
url https://arxiv.org/abs/2510.03093