F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Xiaohui, Xiao, Ruitong, Mo, Jianye, Wu, Bowen, Yu, Qun, Wang, Baoxun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913802120855552
author Sun, Xiaohui
Xiao, Ruitong
Mo, Jianye
Wu, Bowen
Yu, Qun
Wang, Baoxun
author_facet Sun, Xiaohui
Xiao, Ruitong
Mo, Jianye
Wu, Bowen
Yu, Qun
Wang, Baoxun
contents We present F5R-TTS, a novel text-to-speech (TTS) system that integrates Group Relative Policy Optimization (GRPO) into a flow-matching based architecture. By reformulating the deterministic outputs of flow-matching TTS into probabilistic Gaussian distributions, our approach enables seamless integration of reinforcement learning algorithms. During pretraining, we train a probabilistically reformulated flow-matching based model which is derived from F5-TTS with an open-source dataset. In the subsequent reinforcement learning (RL) phase, we employ a GRPO-driven enhancement stage that leverages dual reward metrics: word error rate (WER) computed via automatic speech recognition and speaker similarity (SIM) assessed by verification models. Experimental results on zero-shot voice cloning demonstrate that F5R-TTS achieves significant improvements in both speech intelligibility (a 29.5% relative reduction in WER) and speaker similarity (a 4.6% relative increase in SIM score) compared to conventional flow-matching based TTS systems. Audio samples are available at https://frontierlabs.github.io/F5R.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02407
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
Sun, Xiaohui
Xiao, Ruitong
Mo, Jianye
Wu, Bowen
Yu, Qun
Wang, Baoxun
Sound
Audio and Speech Processing
We present F5R-TTS, a novel text-to-speech (TTS) system that integrates Group Relative Policy Optimization (GRPO) into a flow-matching based architecture. By reformulating the deterministic outputs of flow-matching TTS into probabilistic Gaussian distributions, our approach enables seamless integration of reinforcement learning algorithms. During pretraining, we train a probabilistically reformulated flow-matching based model which is derived from F5-TTS with an open-source dataset. In the subsequent reinforcement learning (RL) phase, we employ a GRPO-driven enhancement stage that leverages dual reward metrics: word error rate (WER) computed via automatic speech recognition and speaker similarity (SIM) assessed by verification models. Experimental results on zero-shot voice cloning demonstrate that F5R-TTS achieves significant improvements in both speech intelligibility (a 29.5% relative reduction in WER) and speaker similarity (a 4.6% relative increase in SIM score) compared to conventional flow-matching based TTS systems. Audio samples are available at https://frontierlabs.github.io/F5R.
title F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2504.02407