DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Hanke, Guo, Dake, Wang, Chengyou, Li, Yue, Tian, Wenjie, Zhu, Xinfa, Wang, Xinsheng, Li, Xiulin, Miao, Guanqiong, Liu, Bo, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917044414316544
author Xie, Hanke
Guo, Dake
Wang, Chengyou
Li, Yue
Tian, Wenjie
Zhu, Xinfa
Wang, Xinsheng
Li, Xiulin
Miao, Guanqiong
Liu, Bo
Xie, Lei
author_facet Xie, Hanke
Guo, Dake
Wang, Chengyou
Li, Yue
Tian, Wenjie
Zhu, Xinfa
Wang, Xinsheng
Li, Xiulin
Miao, Guanqiong
Liu, Bo
Xie, Lei
contents Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However, generating human-like, interactive dialogue speech remains challenging. Current systems face limitations due to the scarcity of dual-track data and difficulties in achieving naturalness, contextual coherence, and interactional dynamics, such as turn-taking, overlapping speech, and speaker consistency, in multi-turn conversations. To address these challenges, we propose DialoSpeech, a dual-track architecture combining a large language model with Chunked Flow Matching for expressive, human-like dialogue speech synthesis. DialoSpeech generates natural multi-turn conversations with coherent speaker turns and natural overlaps, supporting both Chinese and English and cross-lingual speech synthesis. We introduce a data processing pipeline to construct dual-track dialogue datasets, facilitating scalable training and experimental validation. Experiments show that our model outperforms baselines, offering a solution for generating human-like spoken dialogues. Audio samples are available at https://tiamojames.github.io/DialoSpeech
format Preprint
id arxiv_https___arxiv_org_abs_2510_08373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
Xie, Hanke
Guo, Dake
Wang, Chengyou
Li, Yue
Tian, Wenjie
Zhu, Xinfa
Wang, Xinsheng
Li, Xiulin
Miao, Guanqiong
Liu, Bo
Xie, Lei
Audio and Speech Processing
Sound
Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However, generating human-like, interactive dialogue speech remains challenging. Current systems face limitations due to the scarcity of dual-track data and difficulties in achieving naturalness, contextual coherence, and interactional dynamics, such as turn-taking, overlapping speech, and speaker consistency, in multi-turn conversations. To address these challenges, we propose DialoSpeech, a dual-track architecture combining a large language model with Chunked Flow Matching for expressive, human-like dialogue speech synthesis. DialoSpeech generates natural multi-turn conversations with coherent speaker turns and natural overlaps, supporting both Chinese and English and cross-lingual speech synthesis. We introduce a data processing pipeline to construct dual-track dialogue datasets, facilitating scalable training and experimental validation. Experiments show that our model outperforms baselines, offering a solution for generating human-like spoken dialogues. Audio samples are available at https://tiamojames.github.io/DialoSpeech
title DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.08373