PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Jing, Wang, Jiaqi, Tan, Daxin, Chen, Xiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
di: Xu, Jing, et al.
Pubblicazione: (2024)
di: Xu, Jing, et al.
Pubblicazione: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2025)
di: Deng, Keqi, et al.
Pubblicazione: (2025)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
di: Lin, Ju, et al.
Pubblicazione: (2026)
di: Lin, Ju, et al.
Pubblicazione: (2026)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
di: Luu, Nam, et al.
Pubblicazione: (2025)
di: Luu, Nam, et al.
Pubblicazione: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
di: Wang, Jianjin, et al.
Pubblicazione: (2025)
di: Wang, Jianjin, et al.
Pubblicazione: (2025)
Contrastive Feedback Mechanism for Simultaneous Speech Translation
di: Tan, Haotian, et al.
Pubblicazione: (2024)
di: Tan, Haotian, et al.
Pubblicazione: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
di: Koneru, Sai, et al.
Pubblicazione: (2024)
di: Koneru, Sai, et al.
Pubblicazione: (2024)
FASST: Fast LLM-based Simultaneous Speech Translation
di: Ouyang, Siqi, et al.
Pubblicazione: (2024)
di: Ouyang, Siqi, et al.
Pubblicazione: (2024)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
di: Cao, Di, et al.
Pubblicazione: (2026)
di: Cao, Di, et al.
Pubblicazione: (2026)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
di: Liu, Heyang, et al.
Pubblicazione: (2025)
di: Liu, Heyang, et al.
Pubblicazione: (2025)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
di: Zhang, Pei, et al.
Pubblicazione: (2025)
di: Zhang, Pei, et al.
Pubblicazione: (2025)
Zero-resource Speech Translation and Recognition with LLMs
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
di: Pareras, Oriol, et al.
Pubblicazione: (2025)
di: Pareras, Oriol, et al.
Pubblicazione: (2025)
SpeechT: Findings of the First Mentorship in Speech Translation
di: Moslem, Yasmin, et al.
Pubblicazione: (2025)
di: Moslem, Yasmin, et al.
Pubblicazione: (2025)
SpeechQE: Estimating the Quality of Direct Speech Translation
di: Han, HyoJung, et al.
Pubblicazione: (2024)
di: Han, HyoJung, et al.
Pubblicazione: (2024)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
di: Du, Yexing, et al.
Pubblicazione: (2024)
di: Du, Yexing, et al.
Pubblicazione: (2024)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
di: Popescu-Belis, Andrei, et al.
Pubblicazione: (2025)
di: Popescu-Belis, Andrei, et al.
Pubblicazione: (2025)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
di: Alastruey, Belen, et al.
Pubblicazione: (2023)
di: Alastruey, Belen, et al.
Pubblicazione: (2023)
SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
di: Tan, Haotian, et al.
Pubblicazione: (2025)
di: Tan, Haotian, et al.
Pubblicazione: (2025)
Chain-of-Thought Prompting for Speech Translation
di: Hu, Ke, et al.
Pubblicazione: (2024)
di: Hu, Ke, et al.
Pubblicazione: (2024)
SparQLe: Speech Queries to Text Translation Through LLMs
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2025)
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2025)
StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation
di: Chen, Xi, et al.
Pubblicazione: (2025)
di: Chen, Xi, et al.
Pubblicazione: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025)
di: Song, Yuhan, et al.
Pubblicazione: (2025)
SpeechMapper: Speech-to-text Embedding Projector for LLMs
di: Mohapatra, Biswesh, et al.
Pubblicazione: (2026)
di: Mohapatra, Biswesh, et al.
Pubblicazione: (2026)
CMU's IWSLT 2025 Simultaneous Speech Translation System
di: Ouyang, Siqi, et al.
Pubblicazione: (2025)
di: Ouyang, Siqi, et al.
Pubblicazione: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
di: Li, Xuanchen, et al.
Pubblicazione: (2025)
di: Li, Xuanchen, et al.
Pubblicazione: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
MooER: LLM-based Speech Recognition and Translation Models from Moore Threads
di: Xu, Junhao, et al.
Pubblicazione: (2024)
di: Xu, Junhao, et al.
Pubblicazione: (2024)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
di: Xie, Yuan, et al.
Pubblicazione: (2026)
di: Xie, Yuan, et al.
Pubblicazione: (2026)
From TOWER to SPIRE: Adding the Speech Modality to a Translation-Specialist LLM
di: Ambilduke, Kshitij, et al.
Pubblicazione: (2025)
di: Ambilduke, Kshitij, et al.
Pubblicazione: (2025)
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
di: Macháček, Dominik, et al.
Pubblicazione: (2025)
di: Macháček, Dominik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
di: Xu, Jing, et al.
Pubblicazione: (2024) -
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2025) -
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
di: Cui, Wenqian, et al.
Pubblicazione: (2026) -
Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
di: Lin, Ju, et al.
Pubblicazione: (2026) -
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
di: Luu, Nam, et al.
Pubblicazione: (2025)