SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Hanke, Lin, Haopeng, Cao, Wenxiao, Guo, Dake, Tian, Wenjie, Wu, Jun, Wen, Hanlin, Shang, Ruixuan, Liu, Hongmei, Jiang, Zhiqi, Jiang, Yuepeng, Chen, Wenxi, Yan, Ruiqi, Qian, Jiale, Yan, Yichao, Yin, Shunshun, Tao, Ming, Chen, Xie, Xie, Lei, Wang, Xinsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911236021551104
author Xie, Hanke
Lin, Haopeng
Cao, Wenxiao
Guo, Dake
Tian, Wenjie
Wu, Jun
Wen, Hanlin
Shang, Ruixuan
Liu, Hongmei
Jiang, Zhiqi
Jiang, Yuepeng
Chen, Wenxi
Yan, Ruiqi
Qian, Jiale
Yan, Yichao
Yin, Shunshun
Tao, Ming
Chen, Xie
Xie, Lei
Wang, Xinsheng
author_facet Xie, Hanke
Lin, Haopeng
Cao, Wenxiao
Guo, Dake
Tian, Wenjie
Wu, Jun
Wen, Hanlin
Shang, Ruixuan
Liu, Hongmei
Jiang, Zhiqi
Jiang, Yuepeng
Chen, Wenxi
Yan, Ruiqi
Qian, Jiale
Yan, Yichao
Yin, Shunshun
Tao, Ming
Chen, Xie
Xie, Lei
Wang, Xinsheng
contents Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks. To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
Xie, Hanke
Lin, Haopeng
Cao, Wenxiao
Guo, Dake
Tian, Wenjie
Wu, Jun
Wen, Hanlin
Shang, Ruixuan
Liu, Hongmei
Jiang, Zhiqi
Jiang, Yuepeng
Chen, Wenxi
Yan, Ruiqi
Qian, Jiale
Yan, Yichao
Yin, Shunshun
Tao, Ming
Chen, Xie
Xie, Lei
Wang, Xinsheng
Audio and Speech Processing
Sound
Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks. To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis.
title SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.23541