MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jiang, Ziyue, Ren, Yi, Li, Ruiqi, Ji, Shengpeng, Zhang, Boyang, Ye, Zhenhui, Zhang, Chen, Jionghao, Bai, Yang, Xiaoda, Zuo, Jialong, Zhang, Yu, Liu, Rui, Yin, Xiang, Zhao, Zhou
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912298557243392
author Jiang, Ziyue
Ren, Yi
Li, Ruiqi
Ji, Shengpeng
Zhang, Boyang
Ye, Zhenhui
Zhang, Chen
Jionghao, Bai
Yang, Xiaoda
Zuo, Jialong
Zhang, Yu
Liu, Rui
Yin, Xiang
Zhao, Zhou
author_facet Jiang, Ziyue
Ren, Yi
Li, Ruiqi
Ji, Shengpeng
Zhang, Boyang
Ye, Zhenhui
Zhang, Chen
Jionghao, Bai
Yang, Xiaoda
Zuo, Jialong
Zhang, Yu
Liu, Rui
Yin, Xiang
Zhao, Zhou
contents While recent zero-shot text-to-speech (TTS) models have significantly improved speech quality and expressiveness, mainstream systems still suffer from issues related to speech-text alignment modeling: 1) models without explicit speech-text alignment modeling exhibit less robustness, especially for hard sentences in practical applications; 2) predefined alignment-based models suffer from naturalness constraints of forced alignments. This paper introduces \textit{MegaTTS 3}, a TTS system featuring an innovative sparse alignment algorithm that guides the latent diffusion transformer (DiT). Specifically, we provide sparse alignment boundaries to MegaTTS 3 to reduce the difficulty of alignment without limiting the search space, thereby achieving high naturalness. Moreover, we employ a multi-condition classifier-free guidance strategy for accent intensity adjustment and adopt the piecewise rectified flow technique to accelerate the generation process. Experiments demonstrate that MegaTTS 3 achieves state-of-the-art zero-shot TTS speech quality and supports highly flexible control over accent intensity. Notably, our system can generate high-quality one-minute speech with only 8 sampling steps. Audio samples are available at https://sditdemo.github.io/sditdemo/.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Jiang, Ziyue
Ren, Yi
Li, Ruiqi
Ji, Shengpeng
Zhang, Boyang
Ye, Zhenhui
Zhang, Chen
Jionghao, Bai
Yang, Xiaoda
Zuo, Jialong
Zhang, Yu
Liu, Rui
Yin, Xiang
Zhao, Zhou
Audio and Speech Processing
Machine Learning
Sound
While recent zero-shot text-to-speech (TTS) models have significantly improved speech quality and expressiveness, mainstream systems still suffer from issues related to speech-text alignment modeling: 1) models without explicit speech-text alignment modeling exhibit less robustness, especially for hard sentences in practical applications; 2) predefined alignment-based models suffer from naturalness constraints of forced alignments. This paper introduces \textit{MegaTTS 3}, a TTS system featuring an innovative sparse alignment algorithm that guides the latent diffusion transformer (DiT). Specifically, we provide sparse alignment boundaries to MegaTTS 3 to reduce the difficulty of alignment without limiting the search space, thereby achieving high naturalness. Moreover, we employ a multi-condition classifier-free guidance strategy for accent intensity adjustment and adopt the piecewise rectified flow technique to accelerate the generation process. Experiments demonstrate that MegaTTS 3 achieves state-of-the-art zero-shot TTS speech quality and supports highly flexible control over accent intensity. Notably, our system can generate high-quality one-minute speech with only 8 sampling steps. Audio samples are available at https://sditdemo.github.io/sditdemo/.
title MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2502.18924