PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jing, Wang, Jiaqi, Tan, Daxin, Chen, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)
by: Deng, Keqi, et al.
Published: (2025)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
by: Cui, Wenqian, et al.
Published: (2026)
by: Cui, Wenqian, et al.
Published: (2026)
Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
by: Lin, Ju, et al.
Published: (2026)
by: Lin, Ju, et al.
Published: (2026)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025)
by: Luu, Nam, et al.
Published: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
by: Futami, Hayato, et al.
Published: (2025)
by: Futami, Hayato, et al.
Published: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
by: Wang, Jianjin, et al.
Published: (2025)
by: Wang, Jianjin, et al.
Published: (2025)
Contrastive Feedback Mechanism for Simultaneous Speech Translation
by: Tan, Haotian, et al.
Published: (2024)
by: Tan, Haotian, et al.
Published: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
by: Liu, Henglyu, et al.
Published: (2025)
by: Liu, Henglyu, et al.
Published: (2025)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
FASST: Fast LLM-based Simultaneous Speech Translation
by: Ouyang, Siqi, et al.
Published: (2024)
by: Ouyang, Siqi, et al.
Published: (2024)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
by: Cao, Di, et al.
Published: (2026)
by: Cao, Di, et al.
Published: (2026)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
by: Lu, Xiangyu, et al.
Published: (2025)
by: Lu, Xiangyu, et al.
Published: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
by: Parcollet, Titouan, et al.
Published: (2026)
by: Parcollet, Titouan, et al.
Published: (2026)
SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
Zero-resource Speech Translation and Recognition with LLMs
by: Mundnich, Karel, et al.
Published: (2024)
by: Mundnich, Karel, et al.
Published: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
by: Pareras, Oriol, et al.
Published: (2025)
by: Pareras, Oriol, et al.
Published: (2025)
SpeechT: Findings of the First Mentorship in Speech Translation
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
SpeechQE: Estimating the Quality of Direct Speech Translation
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
by: Zhang, Linhao, et al.
Published: (2025)
by: Zhang, Linhao, et al.
Published: (2025)
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
by: Du, Yexing, et al.
Published: (2024)
by: Du, Yexing, et al.
Published: (2024)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
by: Popescu-Belis, Andrei, et al.
Published: (2025)
by: Popescu-Belis, Andrei, et al.
Published: (2025)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
by: Alastruey, Belen, et al.
Published: (2023)
by: Alastruey, Belen, et al.
Published: (2023)
SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
by: Tan, Haotian, et al.
Published: (2025)
by: Tan, Haotian, et al.
Published: (2025)
Chain-of-Thought Prompting for Speech Translation
by: Hu, Ke, et al.
Published: (2024)
by: Hu, Ke, et al.
Published: (2024)
SparQLe: Speech Queries to Text Translation Through LLMs
by: Djanibekov, Amirbek, et al.
Published: (2025)
by: Djanibekov, Amirbek, et al.
Published: (2025)
StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
by: Song, Yuhan, et al.
Published: (2025)
by: Song, Yuhan, et al.
Published: (2025)
SpeechMapper: Speech-to-text Embedding Projector for LLMs
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
CMU's IWSLT 2025 Simultaneous Speech Translation System
by: Ouyang, Siqi, et al.
Published: (2025)
by: Ouyang, Siqi, et al.
Published: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
MooER: LLM-based Speech Recognition and Translation Models from Moore Threads
by: Xu, Junhao, et al.
Published: (2024)
by: Xu, Junhao, et al.
Published: (2024)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
by: Xie, Yuan, et al.
Published: (2026)
by: Xie, Yuan, et al.
Published: (2026)
From TOWER to SPIRE: Adding the Speech Modality to a Translation-Specialist LLM
by: Ambilduke, Kshitij, et al.
Published: (2025)
by: Ambilduke, Kshitij, et al.
Published: (2025)
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
by: Macháček, Dominik, et al.
Published: (2025)
by: Macháček, Dominik, et al.
Published: (2025)
Similar Items
-
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
by: Xu, Jing, et al.
Published: (2024) -
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025) -
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
by: Cui, Wenqian, et al.
Published: (2026) -
Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
by: Lin, Ju, et al.
Published: (2026) -
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025)