Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913961846243328 |
|---|---|
| author | Cheng, Shanbo Bao, Yu Huang, Zhichao Lu, Yu Peng, Ningxin Xu, Lu Yu, Runsheng Cao, Rong Du, Yujiao Han, Ting Hu, Yuxiang Li, Zeyang Liu, Sitong Ma, Shengtao Pan, Shiguang Xiao, Jiongchen Xu, Nuo Yang, Meng Ye, Rong Yu, Yiming Zhang, Jun Zhang, Ruofei Zhang, Wanyi Zhu, Wenhao Zou, Liehao Lu, Lu Wang, Yuxuan Wu, Yonghui |
| author_facet | Cheng, Shanbo Bao, Yu Huang, Zhichao Lu, Yu Peng, Ningxin Xu, Lu Yu, Runsheng Cao, Rong Du, Yujiao Han, Ting Hu, Yuxiang Li, Zeyang Liu, Sitong Ma, Shengtao Pan, Shiguang Xiao, Jiongchen Xu, Nuo Yang, Meng Ye, Rong Yu, Yiming Zhang, Jun Zhang, Ruofei Zhang, Wanyi Zhu, Wenhao Zou, Liehao Lu, Lu Wang, Yuxuan Wu, Yonghui |
| contents | Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_17527 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice Cheng, Shanbo Bao, Yu Huang, Zhichao Lu, Yu Peng, Ningxin Xu, Lu Yu, Runsheng Cao, Rong Du, Yujiao Han, Ting Hu, Yuxiang Li, Zeyang Liu, Sitong Ma, Shengtao Pan, Shiguang Xiao, Jiongchen Xu, Nuo Yang, Meng Ye, Rong Yu, Yiming Zhang, Jun Zhang, Ruofei Zhang, Wanyi Zhu, Wenhao Zou, Liehao Lu, Lu Wang, Yuxuan Wu, Yonghui Computation and Language Sound Audio and Speech Processing Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability. |
| title | Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2507.17527 |