Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Romero-Díaz, Jacobo, Gállego, Gerard I., Pareras, Oriol, Costa, Federico, Hernando, Javier, España-Bonet, Cristina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908574633951232
author Romero-Díaz, Jacobo
Gállego, Gerard I.
Pareras, Oriol
Costa, Federico
Hernando, Javier
España-Bonet, Cristina
author_facet Romero-Díaz, Jacobo
Gállego, Gerard I.
Pareras, Oriol
Costa, Federico
Hernando, Javier
España-Bonet, Cristina
contents Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or other acoustic cues. Chain-of-Thought (CoT) prompting has recently been introduced, with the expectation that jointly accessing speech and transcription will overcome these issues. Analyzing CoT through attribution methods, robustness evaluations with corrupted transcripts, and prosody-awareness, we find that it largely mirrors cascaded behavior, relying mainly on transcripts while barely leveraging speech. Simple training interventions, such as adding Direct S2TT data or noisy transcript injection, enhance robustness and increase speech attribution. These findings challenge the assumed advantages of CoT and highlight the need for architectures that explicitly integrate acoustic information into translation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
Romero-Díaz, Jacobo
Gállego, Gerard I.
Pareras, Oriol
Costa, Federico
Hernando, Javier
España-Bonet, Cristina
Computation and Language
Sound
Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or other acoustic cues. Chain-of-Thought (CoT) prompting has recently been introduced, with the expectation that jointly accessing speech and transcription will overcome these issues. Analyzing CoT through attribution methods, robustness evaluations with corrupted transcripts, and prosody-awareness, we find that it largely mirrors cascaded behavior, relying mainly on transcripts while barely leveraging speech. Simple training interventions, such as adding Direct S2TT data or noisy transcript injection, enhance robustness and increase speech attribution. These findings challenge the assumed advantages of CoT and highlight the need for architectures that explicitly integrate acoustic information into translation.
title Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
topic Computation and Language
Sound
url https://arxiv.org/abs/2510.03115