F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Sun, Xiaohui, Xiao, Ruitong, Mo, Jianye, Wu, Bowen, Yu, Qun, Wang, Baoxun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
por: Chen, Yushen, et al.
Publicado: (2024)
por: Chen, Yushen, et al.
Publicado: (2024)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
por: Glazer, Neta, et al.
Publicado: (2025)
por: Glazer, Neta, et al.
Publicado: (2025)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
por: Guan, Wenhao, et al.
Publicado: (2023)
por: Guan, Wenhao, et al.
Publicado: (2023)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
por: Ren, Yong, et al.
Publicado: (2026)
por: Ren, Yong, et al.
Publicado: (2026)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
por: Li, Xin, et al.
Publicado: (2025)
por: Li, Xin, et al.
Publicado: (2025)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
por: Kim, Jaehyeon, et al.
Publicado: (2024)
por: Kim, Jaehyeon, et al.
Publicado: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
por: Liu, Huadai, et al.
Publicado: (2023)
por: Liu, Huadai, et al.
Publicado: (2023)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
por: Wang, Chunhui, et al.
Publicado: (2024)
por: Wang, Chunhui, et al.
Publicado: (2024)
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
por: Gu, Yi, et al.
Publicado: (2026)
por: Gu, Yi, et al.
Publicado: (2026)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
por: Guo, Hao-Han, et al.
Publicado: (2025)
por: Guo, Hao-Han, et al.
Publicado: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
por: Guo, Hao-Han, et al.
Publicado: (2024)
por: Guo, Hao-Han, et al.
Publicado: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
por: Wu, Zhichao, et al.
Publicado: (2025)
por: Wu, Zhichao, et al.
Publicado: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024)
por: Fu, Ruibo, et al.
Publicado: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
por: Guo, Yinlin, et al.
Publicado: (2024)
por: Guo, Yinlin, et al.
Publicado: (2024)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
por: Guan, Wenhao, et al.
Publicado: (2025)
por: Guan, Wenhao, et al.
Publicado: (2025)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
por: Zheng, Qixi, et al.
Publicado: (2025)
por: Zheng, Qixi, et al.
Publicado: (2025)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
por: Zhu, Han, et al.
Publicado: (2025)
por: Zhu, Han, et al.
Publicado: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
por: Guan, Wenhao, et al.
Publicado: (2023)
por: Guan, Wenhao, et al.
Publicado: (2023)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2025)
por: Lou, Haowei, et al.
Publicado: (2025)
MunTTS: A Text-to-Speech System for Mundari
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
por: Yao, Jixun, et al.
Publicado: (2025)
por: Yao, Jixun, et al.
Publicado: (2025)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
por: Wang, Jianzong, et al.
Publicado: (2023)
por: Wang, Jianzong, et al.
Publicado: (2023)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
por: Bondaruk, Łukasz, et al.
Publicado: (2025)
por: Bondaruk, Łukasz, et al.
Publicado: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
por: Li, Haoxun, et al.
Publicado: (2025)
por: Li, Haoxun, et al.
Publicado: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
por: Xue, Heyang, et al.
Publicado: (2025)
por: Xue, Heyang, et al.
Publicado: (2025)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
por: Lan, Gael Le, et al.
Publicado: (2024)
por: Lan, Gael Le, et al.
Publicado: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
por: Jung, Chaeyoung, et al.
Publicado: (2024)
por: Jung, Chaeyoung, et al.
Publicado: (2024)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
por: Chary, Podakanti Satyajith
Publicado: (2024)
por: Chary, Podakanti Satyajith
Publicado: (2024)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
por: Yang, Dong, et al.
Publicado: (2025)
por: Yang, Dong, et al.
Publicado: (2025)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
por: Tang, Haobin, et al.
Publicado: (2024)
por: Tang, Haobin, et al.
Publicado: (2024)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
por: Kim, Hyeongju, et al.
Publicado: (2025)
por: Kim, Hyeongju, et al.
Publicado: (2025)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2026)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2026)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
por: Das, Shoutrik, et al.
Publicado: (2025)
por: Das, Shoutrik, et al.
Publicado: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2023)
por: Jiang, Ziyue, et al.
Publicado: (2023)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
por: Anastassiou, Philip, et al.
Publicado: (2024)
por: Anastassiou, Philip, et al.
Publicado: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
por: Yang, Da-Hee, et al.
Publicado: (2026)
por: Yang, Da-Hee, et al.
Publicado: (2026)
Ejemplares similares
-
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
por: Chen, Yushen, et al.
Publicado: (2024) -
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
por: Glazer, Neta, et al.
Publicado: (2025) -
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
por: Guan, Wenhao, et al.
Publicado: (2023) -
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
por: Ren, Yong, et al.
Publicado: (2026) -
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
por: Li, Xin, et al.
Publicado: (2025)