Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cheng, Shanbo, Bao, Yu, Huang, Zhichao, Lu, Yu, Peng, Ningxin, Xu, Lu, Yu, Runsheng, Cao, Rong, Du, Yujiao, Han, Ting, Hu, Yuxiang, Li, Zeyang, Liu, Sitong, Ma, Shengtao, Pan, Shiguang, Xiao, Jiongchen, Xu, Nuo, Yang, Meng, Ye, Rong, Yu, Yiming, Zhang, Jun, Zhang, Ruofei, Zhang, Wanyi, Zhu, Wenhao, Zou, Liehao, Lu, Lu, Wang, Yuxuan, Wu, Yonghui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913961846243328
author Cheng, Shanbo
Bao, Yu
Huang, Zhichao
Lu, Yu
Peng, Ningxin
Xu, Lu
Yu, Runsheng
Cao, Rong
Du, Yujiao
Han, Ting
Hu, Yuxiang
Li, Zeyang
Liu, Sitong
Ma, Shengtao
Pan, Shiguang
Xiao, Jiongchen
Xu, Nuo
Yang, Meng
Ye, Rong
Yu, Yiming
Zhang, Jun
Zhang, Ruofei
Zhang, Wanyi
Zhu, Wenhao
Zou, Liehao
Lu, Lu
Wang, Yuxuan
Wu, Yonghui
author_facet Cheng, Shanbo
Bao, Yu
Huang, Zhichao
Lu, Yu
Peng, Ningxin
Xu, Lu
Yu, Runsheng
Cao, Rong
Du, Yujiao
Han, Ting
Hu, Yuxiang
Li, Zeyang
Liu, Sitong
Ma, Shengtao
Pan, Shiguang
Xiao, Jiongchen
Xu, Nuo
Yang, Meng
Ye, Rong
Yu, Yiming
Zhang, Jun
Zhang, Ruofei
Zhang, Wanyi
Zhu, Wenhao
Zou, Liehao
Lu, Lu
Wang, Yuxuan
Wu, Yonghui
contents Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
Cheng, Shanbo
Bao, Yu
Huang, Zhichao
Lu, Yu
Peng, Ningxin
Xu, Lu
Yu, Runsheng
Cao, Rong
Du, Yujiao
Han, Ting
Hu, Yuxiang
Li, Zeyang
Liu, Sitong
Ma, Shengtao
Pan, Shiguang
Xiao, Jiongchen
Xu, Nuo
Yang, Meng
Ye, Rong
Yu, Yiming
Zhang, Jun
Zhang, Ruofei
Zhang, Wanyi
Zhu, Wenhao
Zou, Liehao
Lu, Lu
Wang, Yuxuan
Wu, Yonghui
Computation and Language
Sound
Audio and Speech Processing
Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.
title Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.17527