MoonCast: High-Quality Zero-Shot Podcast Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ju, Zeqian, Yang, Dongchao, Yu, Jianwei, Shen, Kai, Leng, Yichong, Wang, Zhengtao, Tan, Xu, Zhou, Xinyu, Qin, Tao, Li, Xiangyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910883004809216
author Ju, Zeqian
Yang, Dongchao
Yu, Jianwei
Shen, Kai
Leng, Yichong
Wang, Zhengtao
Tan, Xu
Zhou, Xinyu
Qin, Tao
Li, Xiangyang
author_facet Ju, Zeqian
Yang, Dongchao
Yu, Jianwei
Shen, Kai
Leng, Yichong
Wang, Zhengtao
Tan, Xu
Zhou, Xinyu
Qin, Tao
Li, Xiangyang
contents Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scenarios such as podcasts. These limitations arise from two primary challenges: 1) long speech: podcasts typically span several minutes, exceeding the upper limit of most existing work; 2) spontaneity: podcasts are marked by their spontaneous, oral nature, which sharply contrasts with formal, written contexts; existing works often fall short in capturing this spontaneity. In this paper, we propose MoonCast, a solution for high-quality zero-shot podcast generation, aiming to synthesize natural podcast-style speech from text-only sources (e.g., stories, technical reports, news in TXT, PDF, or Web URL formats) using the voices of unseen speakers. To generate long audio, we adopt a long-context language model-based audio modeling approach utilizing large-scale long-context speech data. To enhance spontaneity, we utilize a podcast generation module to generate scripts with spontaneous details, which have been empirically shown to be as crucial as the text-to-speech modeling itself. Experiments demonstrate that MoonCast outperforms baselines, with particularly notable improvements in spontaneity and coherence.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14345
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoonCast: High-Quality Zero-Shot Podcast Generation
Ju, Zeqian
Yang, Dongchao
Yu, Jianwei
Shen, Kai
Leng, Yichong
Wang, Zhengtao
Tan, Xu
Zhou, Xinyu
Qin, Tao
Li, Xiangyang
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scenarios such as podcasts. These limitations arise from two primary challenges: 1) long speech: podcasts typically span several minutes, exceeding the upper limit of most existing work; 2) spontaneity: podcasts are marked by their spontaneous, oral nature, which sharply contrasts with formal, written contexts; existing works often fall short in capturing this spontaneity. In this paper, we propose MoonCast, a solution for high-quality zero-shot podcast generation, aiming to synthesize natural podcast-style speech from text-only sources (e.g., stories, technical reports, news in TXT, PDF, or Web URL formats) using the voices of unseen speakers. To generate long audio, we adopt a long-context language model-based audio modeling approach utilizing large-scale long-context speech data. To enhance spontaneity, we utilize a podcast generation module to generate scripts with spontaneous details, which have been empirically shown to be as crucial as the text-to-speech modeling itself. Experiments demonstrate that MoonCast outperforms baselines, with particularly notable improvements in spontaneity and coherence.
title MoonCast: High-Quality Zero-Shot Podcast Generation
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2503.14345