Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
Fuente:
arXiv
Saved in:
| Main Authors: | Pareras, Oriol, Gállego, Gerard I., Costa, Federico, España-Bonet, Cristina, Hernando, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
by: Gállego, Gerard I., et al.
Published: (2025)
by: Gállego, Gerard I., et al.
Published: (2025)
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
by: Buitrago, Pol, et al.
Published: (2026)
by: Buitrago, Pol, et al.
Published: (2026)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
by: Futami, Hayato, et al.
Published: (2025)
by: Futami, Hayato, et al.
Published: (2025)
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
by: Costa, Federico, et al.
Published: (2024)
by: Costa, Federico, et al.
Published: (2024)
Direct Speech to Speech Translation: A Review
by: Sarim, Mohammad, et al.
Published: (2025)
by: Sarim, Mohammad, et al.
Published: (2025)
Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
by: Xue, Hongfei, et al.
Published: (2025)
by: Xue, Hongfei, et al.
Published: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
by: Shankar, Bhavani, et al.
Published: (2024)
by: Shankar, Bhavani, et al.
Published: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
by: Gupta, Mahendra, et al.
Published: (2024)
by: Gupta, Mahendra, et al.
Published: (2024)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
by: Zhao, Mengjie, et al.
Published: (2026)
by: Zhao, Mengjie, et al.
Published: (2026)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)
by: Deng, Keqi, et al.
Published: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
by: Liu, Henglyu, et al.
Published: (2025)
by: Liu, Henglyu, et al.
Published: (2025)
SpeechT: Findings of the First Mentorship in Speech Translation
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
Scaling Rich Style-Prompted Text-to-Speech Datasets
by: Diwan, Anuj, et al.
Published: (2025)
by: Diwan, Anuj, et al.
Published: (2025)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
by: Zhou, Rui, et al.
Published: (2024)
by: Zhou, Rui, et al.
Published: (2024)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
by: Yang, Fei, et al.
Published: (2026)
by: Yang, Fei, et al.
Published: (2026)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
by: Yin, Kang, et al.
Published: (2025)
by: Yin, Kang, et al.
Published: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data
by: Buitrago, Pol, et al.
Published: (2026)
by: Buitrago, Pol, et al.
Published: (2026)
Enhancing Crowdsourced Audio for Text-to-Speech Models
by: Giraldo, José, et al.
Published: (2024)
by: Giraldo, José, et al.
Published: (2024)
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
by: Magoshi, Ryo, et al.
Published: (2026)
by: Magoshi, Ryo, et al.
Published: (2026)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
by: Adetiba, Emmanuel, et al.
Published: (2025)
by: Adetiba, Emmanuel, et al.
Published: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
by: Lei, Shun, et al.
Published: (2023)
by: Lei, Shun, et al.
Published: (2023)
End-to-End Speech-to-Text Translation: A Survey
by: Sethiya, Nivedita, et al.
Published: (2023)
by: Sethiya, Nivedita, et al.
Published: (2023)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
by: Bott, Thomas, et al.
Published: (2024)
by: Bott, Thomas, et al.
Published: (2024)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
by: Cheng, Sitong, et al.
Published: (2025)
by: Cheng, Sitong, et al.
Published: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text Directions
by: Khadse, Parth, et al.
Published: (2026)
by: Khadse, Parth, et al.
Published: (2026)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
by: Labiausse, Tom, et al.
Published: (2025)
by: Labiausse, Tom, et al.
Published: (2025)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
by: Li, Hanzhao, et al.
Published: (2025)
by: Li, Hanzhao, et al.
Published: (2025)
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
by: Maltais, Marie, et al.
Published: (2026)
by: Maltais, Marie, et al.
Published: (2026)
Bypassing Direct Reconstruction: Speech Detection from MEG via Large-Scale Audio Retrieval
by: Xiao, Boda, et al.
Published: (2026)
by: Xiao, Boda, et al.
Published: (2026)
Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
by: Kim, Nam-Gyu
Published: (2025)
by: Kim, Nam-Gyu
Published: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
by: Goncalves, Lucas, et al.
Published: (2024)
by: Goncalves, Lucas, et al.
Published: (2024)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
Continuous Speech Tokenizer in Text To Speech
by: Li, Yixing, et al.
Published: (2024)
by: Li, Yixing, et al.
Published: (2024)
Similar Items
-
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
by: Romero-Díaz, Jacobo, et al.
Published: (2025) -
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
by: Gállego, Gerard I., et al.
Published: (2025) -
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
by: Buitrago, Pol, et al.
Published: (2026) -
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025) -
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
by: Futami, Hayato, et al.
Published: (2025)