Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Jingyi, Byun, Ju Seung, Elsner, Micha, Wang, Pichao, Perrault, Andrew
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916882036031488
author Chen, Jingyi
Byun, Ju Seung
Elsner, Micha
Wang, Pichao
Perrault, Andrew
author_facet Chen, Jingyi
Byun, Ju Seung
Elsner, Micha
Wang, Pichao
Perrault, Andrew
contents Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO), an RLHF framework for TTS diffusion models. DLPO integrates the original training loss into the reward function, preserving generative capabilities while reducing inefficiencies. Using naturalness scores as feedback, DLPO aligns reward optimization with the diffusion model's structure, improving speech quality. We evaluate DLPO on WaveGrad 2, a non-autoregressive diffusion-based TTS model. Results show significant improvements in objective metrics (UTMOS 3.65, NISQA 4.02) and subjective evaluations, with DLPO audio preferred 67\% of the time. These findings demonstrate DLPO's potential for efficient, high-quality diffusion TTS in real-time, resource-limited settings.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03123
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
Chen, Jingyi
Byun, Ju Seung
Elsner, Micha
Wang, Pichao
Perrault, Andrew
Sound
Artificial Intelligence
Audio and Speech Processing
Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO), an RLHF framework for TTS diffusion models. DLPO integrates the original training loss into the reward function, preserving generative capabilities while reducing inefficiencies. Using naturalness scores as feedback, DLPO aligns reward optimization with the diffusion model's structure, improving speech quality. We evaluate DLPO on WaveGrad 2, a non-autoregressive diffusion-based TTS model. Results show significant improvements in objective metrics (UTMOS 3.65, NISQA 4.02) and subjective evaluations, with DLPO audio preferred 67\% of the time. These findings demonstrate DLPO's potential for efficient, high-quality diffusion TTS in real-time, resource-limited settings.
title Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2508.03123