MOSS-TTSD: Text to Spoken Dialogue Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yuqian, Yu, Donghua, Lin, Zhengyuan, Jiang, Botian, Chen, Mingshu, Jiang, Yaozhou, Zhao, Yiwei, Zhang, Yiyang, Yuan, Yucheng, Chen, Hanfu, Huang, Kexin, Zhan, Jun, Chang, Cheng, Fei, Zhaoye, Li, Shimin, Yang, Xiaogui, Cheng, Qinyuan, Qiu, Xipeng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911531334107136
author Zhang, Yuqian
Yu, Donghua
Lin, Zhengyuan
Jiang, Botian
Chen, Mingshu
Jiang, Yaozhou
Zhao, Yiwei
Zhang, Yiyang
Yuan, Yucheng
Chen, Hanfu
Huang, Kexin
Zhan, Jun
Chang, Cheng
Fei, Zhaoye
Li, Shimin
Yang, Xiaogui
Cheng, Qinyuan
Qiu, Xipeng
author_facet Zhang, Yuqian
Yu, Donghua
Lin, Zhengyuan
Jiang, Botian
Chen, Mingshu
Jiang, Yaozhou
Zhao, Yiwei
Zhang, Yiyang
Yuan, Yucheng
Chen, Hanfu
Huang, Kexin
Zhan, Jun
Chang, Cheng
Fei, Zhaoye
Li, Shimin
Yang, Xiaogui
Cheng, Qinyuan
Qiu, Xipeng
contents Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate turn-taking, cross-turn acoustic consistency, and long-form stability, which current models often fail to address due to a lack of dialogue context modeling. To bridge this gap, we present MOSS-TTSD, a spoken dialogue synthesis model designed for expressive, multi-party conversational speech across multiple languages. With enhanced long-context modeling, MOSS-TTSD generates long-form spoken conversations from dialogue scripts with explicit speaker tags, supporting up to 60 minutes of single-pass synthesis, multi-party dialogue with up to 5 speakers, and zero-shot voice cloning from a short reference audio clip. The model supports various mainstream languages, including English and Chinese, and is adapted to several long-form scenarios. Additionally, to address limitations of existing evaluation methods, we propose TTSD-eval, an objective evaluation framework based on forced alignment that measures speaker attribution accuracy and speaker similarity without relying on speaker diarization tools. Both objective and subjective evaluation results show that MOSS-TTSD surpasses strong open-source and proprietary baselines in dialogue synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MOSS-TTSD: Text to Spoken Dialogue Generation
Zhang, Yuqian
Yu, Donghua
Lin, Zhengyuan
Jiang, Botian
Chen, Mingshu
Jiang, Yaozhou
Zhao, Yiwei
Zhang, Yiyang
Yuan, Yucheng
Chen, Hanfu
Huang, Kexin
Zhan, Jun
Chang, Cheng
Fei, Zhaoye
Li, Shimin
Yang, Xiaogui
Cheng, Qinyuan
Qiu, Xipeng
Sound
Artificial Intelligence
Computation and Language
Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate turn-taking, cross-turn acoustic consistency, and long-form stability, which current models often fail to address due to a lack of dialogue context modeling. To bridge this gap, we present MOSS-TTSD, a spoken dialogue synthesis model designed for expressive, multi-party conversational speech across multiple languages. With enhanced long-context modeling, MOSS-TTSD generates long-form spoken conversations from dialogue scripts with explicit speaker tags, supporting up to 60 minutes of single-pass synthesis, multi-party dialogue with up to 5 speakers, and zero-shot voice cloning from a short reference audio clip. The model supports various mainstream languages, including English and Chinese, and is adapted to several long-form scenarios. Additionally, to address limitations of existing evaluation methods, we propose TTSD-eval, an objective evaluation framework based on forced alignment that measures speaker attribution accuracy and speaker similarity without relying on speaker diarization tools. Both objective and subjective evaluation results show that MOSS-TTSD surpasses strong open-source and proprietary baselines in dialogue synthesis.
title MOSS-TTSD: Text to Spoken Dialogue Generation
topic Sound
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.19739