SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911236021551104 |
|---|---|
| author | Xie, Hanke Lin, Haopeng Cao, Wenxiao Guo, Dake Tian, Wenjie Wu, Jun Wen, Hanlin Shang, Ruixuan Liu, Hongmei Jiang, Zhiqi Jiang, Yuepeng Chen, Wenxi Yan, Ruiqi Qian, Jiale Yan, Yichao Yin, Shunshun Tao, Ming Chen, Xie Xie, Lei Wang, Xinsheng |
| author_facet | Xie, Hanke Lin, Haopeng Cao, Wenxiao Guo, Dake Tian, Wenjie Wu, Jun Wen, Hanlin Shang, Ruixuan Liu, Hongmei Jiang, Zhiqi Jiang, Yuepeng Chen, Wenxi Yan, Ruiqi Qian, Jiale Yan, Yichao Yin, Shunshun Tao, Ming Chen, Xie Xie, Lei Wang, Xinsheng |
| contents | Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks.
To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23541 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity Xie, Hanke Lin, Haopeng Cao, Wenxiao Guo, Dake Tian, Wenjie Wu, Jun Wen, Hanlin Shang, Ruixuan Liu, Hongmei Jiang, Zhiqi Jiang, Yuepeng Chen, Wenxi Yan, Ruiqi Qian, Jiale Yan, Yichao Yin, Shunshun Tao, Ming Chen, Xie Xie, Lei Wang, Xinsheng Audio and Speech Processing Sound Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks. To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis. |
| title | SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2510.23541 |