Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.19090 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912782060879872 |
|---|---|
| author | Yu, Fan Wang, Tao Wu, You Zhu, Lin Deng, Wei Han, Weisheng Wang, Wenchao Hu, Lin Liang, Xiangyu He, Xiaodong Huang, Yankun Gu, Yu Liu, Yuan Wang, Yuxuan Xiao, Zhangyu Wang, Ziteng Dong, Boya Dang, Feng Chen, Jinming Li, Jingdong Wang, Jun Jin, Yechen Zhang, Yuan Sheng, Zhengyan Wang, Xin |
| author_facet | Yu, Fan Wang, Tao Wu, You Zhu, Lin Deng, Wei Han, Weisheng Wang, Wenchao Hu, Lin Liang, Xiangyu He, Xiaodong Huang, Yankun Gu, Yu Liu, Yuan Wang, Yuxuan Xiao, Zhangyu Wang, Ziteng Dong, Boya Dang, Feng Chen, Jinming Li, Jingdong Wang, Jun Jin, Yechen Zhang, Yuan Sheng, Zhengyan Wang, Xin |
| contents | Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_19090 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis Yu, Fan Wang, Tao Wu, You Zhu, Lin Deng, Wei Han, Weisheng Wang, Wenchao Hu, Lin Liang, Xiangyu He, Xiaodong Huang, Yankun Gu, Yu Liu, Yuan Wang, Yuxuan Xiao, Zhangyu Wang, Ziteng Dong, Boya Dang, Feng Chen, Jinming Li, Jingdong Wang, Jun Jin, Yechen Zhang, Yuan Sheng, Zhengyan Wang, Xin Sound Audio and Speech Processing Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice |
| title | JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.19090 |