F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913802120855552 |
|---|---|
| author | Sun, Xiaohui Xiao, Ruitong Mo, Jianye Wu, Bowen Yu, Qun Wang, Baoxun |
| author_facet | Sun, Xiaohui Xiao, Ruitong Mo, Jianye Wu, Bowen Yu, Qun Wang, Baoxun |
| contents | We present F5R-TTS, a novel text-to-speech (TTS) system that integrates Group Relative Policy Optimization (GRPO) into a flow-matching based architecture. By reformulating the deterministic outputs of flow-matching TTS into probabilistic Gaussian distributions, our approach enables seamless integration of reinforcement learning algorithms. During pretraining, we train a probabilistically reformulated flow-matching based model which is derived from F5-TTS with an open-source dataset. In the subsequent reinforcement learning (RL) phase, we employ a GRPO-driven enhancement stage that leverages dual reward metrics: word error rate (WER) computed via automatic speech recognition and speaker similarity (SIM) assessed by verification models. Experimental results on zero-shot voice cloning demonstrate that F5R-TTS achieves significant improvements in both speech intelligibility (a 29.5% relative reduction in WER) and speaker similarity (a 4.6% relative increase in SIM score) compared to conventional flow-matching based TTS systems. Audio samples are available at https://frontierlabs.github.io/F5R. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_02407 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization Sun, Xiaohui Xiao, Ruitong Mo, Jianye Wu, Bowen Yu, Qun Wang, Baoxun Sound Audio and Speech Processing We present F5R-TTS, a novel text-to-speech (TTS) system that integrates Group Relative Policy Optimization (GRPO) into a flow-matching based architecture. By reformulating the deterministic outputs of flow-matching TTS into probabilistic Gaussian distributions, our approach enables seamless integration of reinforcement learning algorithms. During pretraining, we train a probabilistically reformulated flow-matching based model which is derived from F5-TTS with an open-source dataset. In the subsequent reinforcement learning (RL) phase, we employ a GRPO-driven enhancement stage that leverages dual reward metrics: word error rate (WER) computed via automatic speech recognition and speaker similarity (SIM) assessed by verification models. Experimental results on zero-shot voice cloning demonstrate that F5R-TTS achieves significant improvements in both speech intelligibility (a 29.5% relative reduction in WER) and speaker similarity (a 4.6% relative increase in SIM score) compared to conventional flow-matching based TTS systems. Audio samples are available at https://frontierlabs.github.io/F5R. |
| title | F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2504.02407 |