Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pareras, Oriol, Gállego, Gerard I., Costa, Federico, España-Bonet, Cristina, Hernando, Javier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
von: Costa, Federico, et al.
Veröffentlicht: (2024)
von: Costa, Federico, et al.
Veröffentlicht: (2024)
Direct Speech to Speech Translation: A Review
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
von: Shankar, Bhavani, et al.
Veröffentlicht: (2024)
von: Shankar, Bhavani, et al.
Veröffentlicht: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
SpeechT: Findings of the First Mentorship in Speech Translation
von: Moslem, Yasmin, et al.
Veröffentlicht: (2025)
von: Moslem, Yasmin, et al.
Veröffentlicht: (2025)
Scaling Rich Style-Prompted Text-to-Speech Datasets
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
von: Yang, Fei, et al.
Veröffentlicht: (2026)
von: Yang, Fei, et al.
Veröffentlicht: (2026)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
von: Yin, Kang, et al.
Veröffentlicht: (2025)
von: Yin, Kang, et al.
Veröffentlicht: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024)
von: Giraldo, José, et al.
Veröffentlicht: (2024)
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
von: Magoshi, Ryo, et al.
Veröffentlicht: (2026)
von: Magoshi, Ryo, et al.
Veröffentlicht: (2026)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025)
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text Directions
von: Khadse, Parth, et al.
Veröffentlicht: (2026)
von: Khadse, Parth, et al.
Veröffentlicht: (2026)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
von: Maltais, Marie, et al.
Veröffentlicht: (2026)
von: Maltais, Marie, et al.
Veröffentlicht: (2026)
Bypassing Direct Reconstruction: Speech Detection from MEG via Large-Scale Audio Retrieval
von: Xiao, Boda, et al.
Veröffentlicht: (2026)
von: Xiao, Boda, et al.
Veröffentlicht: (2026)
Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu
Veröffentlicht: (2025)
von: Kim, Nam-Gyu
Veröffentlicht: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025) -
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025) -
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
von: Buitrago, Pol, et al.
Veröffentlicht: (2026) -
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
von: Papi, Sara, et al.
Veröffentlicht: (2025) -
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)