CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Xiang, Bu, Fan, Mehrish, Ambuj, Li, Yingting, Han, Jiale, Cheng, Bo, Poria, Soujanya |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
par: Zhang, Shaozuo, et autres
Publié: (2025)
par: Zhang, Shaozuo, et autres
Publié: (2025)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
par: Li, Yingting, et autres
Publié: (2024)
par: Li, Yingting, et autres
Publié: (2024)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
par: Li, Yingting, et autres
Publié: (2024)
par: Li, Yingting, et autres
Publié: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
par: Melechovsky, Jan, et autres
Publié: (2022)
par: Melechovsky, Jan, et autres
Publié: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
par: Melechovsky, Jan, et autres
Publié: (2024)
par: Melechovsky, Jan, et autres
Publié: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
par: Melechovsky, Jan, et autres
Publié: (2024)
par: Melechovsky, Jan, et autres
Publié: (2024)
Improving Text-To-Audio Models with Synthetic Captions
par: Kong, Zhifeng, et autres
Publié: (2024)
par: Kong, Zhifeng, et autres
Publié: (2024)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
par: Hung, Chia-Yu, et autres
Publié: (2024)
par: Hung, Chia-Yu, et autres
Publié: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
par: Li, Haoxun, et autres
Publié: (2025)
par: Li, Haoxun, et autres
Publié: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
par: Lu, Ye-Xin, et autres
Publié: (2025)
par: Lu, Ye-Xin, et autres
Publié: (2025)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
par: Guo, Yinlin, et autres
Publié: (2024)
par: Guo, Yinlin, et autres
Publié: (2024)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
par: Li, Yinghao Aaron, et autres
Publié: (2024)
par: Li, Yinghao Aaron, et autres
Publié: (2024)
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
par: Kang, Jaeyong, et autres
Publié: (2023)
par: Kang, Jaeyong, et autres
Publié: (2023)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
par: Liu, Huadai, et autres
Publié: (2023)
par: Liu, Huadai, et autres
Publié: (2023)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
par: Jiang, Ziyue, et autres
Publié: (2023)
par: Jiang, Ziyue, et autres
Publié: (2023)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
par: Wang, Jianzong, et autres
Publié: (2023)
par: Wang, Jianzong, et autres
Publié: (2023)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
par: Guo, Hao-Han, et autres
Publié: (2025)
par: Guo, Hao-Han, et autres
Publié: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
par: Guo, Hao-Han, et autres
Publié: (2024)
par: Guo, Hao-Han, et autres
Publié: (2024)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
par: Li, Xin, et autres
Publié: (2025)
par: Li, Xin, et autres
Publié: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
par: Fu, Ruibo, et autres
Publié: (2024)
par: Fu, Ruibo, et autres
Publié: (2024)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
par: Chen, Zhiyong, et autres
Publié: (2024)
par: Chen, Zhiyong, et autres
Publié: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
par: Gong, Cheng, et autres
Publié: (2023)
par: Gong, Cheng, et autres
Publié: (2023)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
par: Kim, Jaehyeon, et autres
Publié: (2024)
par: Kim, Jaehyeon, et autres
Publié: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
par: Xue, Heyang, et autres
Publié: (2025)
par: Xue, Heyang, et autres
Publié: (2025)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
par: Han, Wooseok, et autres
Publié: (2024)
par: Han, Wooseok, et autres
Publié: (2024)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
par: Lou, Haowei, et autres
Publié: (2025)
par: Lou, Haowei, et autres
Publié: (2025)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
par: Zhang, Jiawei, et autres
Publié: (2024)
par: Zhang, Jiawei, et autres
Publié: (2024)
MunTTS: A Text-to-Speech System for Mundari
par: Gumma, Varun, et autres
Publié: (2024)
par: Gumma, Varun, et autres
Publié: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
par: Anastassiou, Philip, et autres
Publié: (2024)
par: Anastassiou, Philip, et autres
Publié: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
par: Li, Hanzhao, et autres
Publié: (2023)
par: Li, Hanzhao, et autres
Publié: (2023)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
par: Wang, Chunhui, et autres
Publié: (2024)
par: Wang, Chunhui, et autres
Publié: (2024)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
par: Sun, Xiaohui, et autres
Publié: (2025)
par: Sun, Xiaohui, et autres
Publié: (2025)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
par: Xie, Kun, et autres
Publié: (2025)
par: Xie, Kun, et autres
Publié: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
par: Liu, Zhijun, et autres
Publié: (2024)
par: Liu, Zhijun, et autres
Publié: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
par: Li, Haitao, et autres
Publié: (2026)
par: Li, Haitao, et autres
Publié: (2026)
Documents similaires
-
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
par: Zhang, Shaozuo, et autres
Publié: (2025) -
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
par: Li, Yingting, et autres
Publié: (2024) -
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
par: Li, Yingting, et autres
Publié: (2024) -
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
par: Melechovsky, Jan, et autres
Publié: (2022) -
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
par: Melechovsky, Jan, et autres
Publié: (2024)