StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shaolei, Fang, Qingkai, Guo, Shoutao, Ma, Zhengrui, Zhang, Min, Feng, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
von: Ma, Zhengrui, et al.
Veröffentlicht: (2025)
von: Ma, Zhengrui, et al.
Veröffentlicht: (2025)
Simultaneous Speech-to-Speech Translation Without Aligned Data
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
von: Yang, Runyan, et al.
Veröffentlicht: (2024)
von: Yang, Runyan, et al.
Veröffentlicht: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
DiariST: Streaming Speech Translation with Speaker Diarization
von: Yang, Mu, et al.
Veröffentlicht: (2023)
von: Yang, Mu, et al.
Veröffentlicht: (2023)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
NAIST Simultaneous Speech Translation System for IWSLT 2024
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
SimulTron: On-Device Simultaneous Speech to Speech Translation
von: Agranovich, Alex, et al.
Veröffentlicht: (2024)
von: Agranovich, Alex, et al.
Veröffentlicht: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Direct Speech to Speech Translation: A Review
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
Direct Speech-to-Speech Neural Machine Translation: A Survey
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024) -
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
von: Fang, Qingkai, et al.
Veröffentlicht: (2024) -
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
von: Fang, Qingkai, et al.
Veröffentlicht: (2025) -
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025) -
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)