Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hussain, Shehzeen, Neekhara, Paarth, Yang, Xuesong, Casanova, Edresson, Ghosh, Subhankar, Desta, Mikyas T., Fejgin, Roy, Valle, Rafael, Li, Jason |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
von: Langman, Ryan, et al.
Veröffentlicht: (2025)
von: Langman, Ryan, et al.
Veröffentlicht: (2025)
Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
von: Fejgin, Roy, et al.
Veröffentlicht: (2025)
von: Fejgin, Roy, et al.
Veröffentlicht: (2025)
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
von: Neekhara, Paarth, et al.
Veröffentlicht: (2024)
von: Neekhara, Paarth, et al.
Veröffentlicht: (2024)
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
von: Neekhara, Paarth, et al.
Veröffentlicht: (2023)
von: Neekhara, Paarth, et al.
Veröffentlicht: (2023)
Towards Flow-Matching-based TTS without Classifier-Free Guidance
von: Liang, Yuzhe, et al.
Veröffentlicht: (2025)
von: Liang, Yuzhe, et al.
Veröffentlicht: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Coherence-Based Frequency Subset Selection For Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
von: Fejgin, Daniel, et al.
Veröffentlicht: (2022)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2022)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting a Calibrated External Microphone Array
von: Fejgin, Daniel, et al.
Veröffentlicht: (2022)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2022)
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2025)
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2025)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
von: Li, Zirui, et al.
Veröffentlicht: (2025)
von: Li, Zirui, et al.
Veröffentlicht: (2025)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
von: Singh, Jaskaran, et al.
Veröffentlicht: (2025)
von: Singh, Jaskaran, et al.
Veröffentlicht: (2025)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
von: Fejgin, Daniel, et al.
Veröffentlicht: (2023)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2023)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
von: Fejgin, Daniel, et al.
Veröffentlicht: (2025)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2025)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
von: Kim, Semin, et al.
Veröffentlicht: (2026)
von: Kim, Semin, et al.
Veröffentlicht: (2026)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
von: Meng, Qingliang, et al.
Veröffentlicht: (2025)
von: Meng, Qingliang, et al.
Veröffentlicht: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
von: Nie, Sihang, et al.
Veröffentlicht: (2025)
von: Nie, Sihang, et al.
Veröffentlicht: (2025)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
BRUDEX Database: Binaural Room Impulse Responses with Uniformly Distributed External Microphones
von: Fejgin, Daniel, et al.
Veröffentlicht: (2023)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2023)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
von: Răgman, Teodora, et al.
Veröffentlicht: (2026)
von: Răgman, Teodora, et al.
Veröffentlicht: (2026)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2025)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025) -
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
von: Langman, Ryan, et al.
Veröffentlicht: (2025) -
Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
von: Fejgin, Roy, et al.
Veröffentlicht: (2025) -
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2025) -
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)