ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Tao, Yu, Yinfeng, Wang, Liejun, Sun, Fuchun, Zheng, Wendong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918155891245056
author Zhu, Tao
Yu, Yinfeng
Wang, Liejun
Sun, Fuchun
Zheng, Wendong
author_facet Zhu, Tao
Yu, Yinfeng
Wang, Liejun
Sun, Fuchun
Zheng, Wendong
contents Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser's ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model's training cost and complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
Zhu, Tao
Yu, Yinfeng
Wang, Liejun
Sun, Fuchun
Zheng, Wendong
Sound
Artificial Intelligence
Audio and Speech Processing
Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser's ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model's training cost and complexity.
title ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.05984