Saved in:
Bibliographic Details
Main Authors: Yu, Fan, Wang, Tao, Wu, You, Zhu, Lin, Deng, Wei, Han, Weisheng, Wang, Wenchao, Hu, Lin, Liang, Xiangyu, He, Xiaodong, Huang, Yankun, Gu, Yu, Liu, Yuan, Wang, Yuxuan, Xiao, Zhangyu, Wang, Ziteng, Dong, Boya, Dang, Feng, Chen, Jinming, Li, Jingdong, Wang, Jun, Jin, Yechen, Zhang, Yuan, Sheng, Zhengyan, Wang, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.19090
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912782060879872
author Yu, Fan
Wang, Tao
Wu, You
Zhu, Lin
Deng, Wei
Han, Weisheng
Wang, Wenchao
Hu, Lin
Liang, Xiangyu
He, Xiaodong
Huang, Yankun
Gu, Yu
Liu, Yuan
Wang, Yuxuan
Xiao, Zhangyu
Wang, Ziteng
Dong, Boya
Dang, Feng
Chen, Jinming
Li, Jingdong
Wang, Jun
Jin, Yechen
Zhang, Yuan
Sheng, Zhengyan
Wang, Xin
author_facet Yu, Fan
Wang, Tao
Wu, You
Zhu, Lin
Deng, Wei
Han, Weisheng
Wang, Wenchao
Hu, Lin
Liang, Xiangyu
He, Xiaodong
Huang, Yankun
Gu, Yu
Liu, Yuan
Wang, Yuxuan
Xiao, Zhangyu
Wang, Ziteng
Dong, Boya
Dang, Feng
Chen, Jinming
Li, Jingdong
Wang, Jun
Jin, Yechen
Zhang, Yuan
Sheng, Zhengyan
Wang, Xin
contents Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice
format Preprint
id arxiv_https___arxiv_org_abs_2512_19090
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
Yu, Fan
Wang, Tao
Wu, You
Zhu, Lin
Deng, Wei
Han, Weisheng
Wang, Wenchao
Hu, Lin
Liang, Xiangyu
He, Xiaodong
Huang, Yankun
Gu, Yu
Liu, Yuan
Wang, Yuxuan
Xiao, Zhangyu
Wang, Ziteng
Dong, Boya
Dang, Feng
Chen, Jinming
Li, Jingdong
Wang, Jun
Jin, Yechen
Zhang, Yuan
Sheng, Zhengyan
Wang, Xin
Sound
Audio and Speech Processing
Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice
title JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2512.19090