RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Cong, Gao, Changfeng, Xiang, Yang, Du, Zhihao, An, Keyu, Zhao, Han, Chen, Qian, Li, Xiangang, Gao, Yingming, Li, Ya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explore the Reinforcement Learning for the LLM based ASR and TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
Differentiable Reward Optimization for LLM based TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
von: An, Keyu, et al.
Veröffentlicht: (2025)
von: An, Keyu, et al.
Veröffentlicht: (2025)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
von: Zhou, Fangru, et al.
Veröffentlicht: (2025)
von: Zhou, Fangru, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
von: Du, Zhihao, et al.
Veröffentlicht: (2025)
von: Du, Zhihao, et al.
Veröffentlicht: (2025)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Explore the Reinforcement Learning for the LLM based ASR and TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025) -
Differentiable Reward Optimization for LLM based TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025) -
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024) -
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
von: An, Keyu, et al.
Veröffentlicht: (2025) -
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)