Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Ziyue, Liu, Jinglin, Ren, Yi, He, Jinzheng, Ye, Zhenhui, Ji, Shengpeng, Yang, Qian, Zhang, Chen, Wei, Pengfei, Wang, Chunfeng, Yin, Xiang, Ma, Zejun, Zhao, Zhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2025)
di: Jiang, Ziyue, et al.
Pubblicazione: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
di: Huang, Jiawei, et al.
Pubblicazione: (2024)
di: Huang, Jiawei, et al.
Pubblicazione: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
di: Liu, Huadai, et al.
Pubblicazione: (2023)
di: Liu, Huadai, et al.
Pubblicazione: (2023)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
di: Li, Haitao, et al.
Pubblicazione: (2026)
di: Li, Haitao, et al.
Pubblicazione: (2026)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
Intelli-Z: Toward Intelligible Zero-Shot TTS
di: Jung, Sunghee, et al.
Pubblicazione: (2024)
di: Jung, Sunghee, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
di: Lei, Shun, et al.
Pubblicazione: (2023)
di: Lei, Shun, et al.
Pubblicazione: (2023)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
di: Deng, Wei, et al.
Pubblicazione: (2025)
di: Deng, Wei, et al.
Pubblicazione: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
di: Dai, Shuqi, et al.
Pubblicazione: (2025)
di: Dai, Shuqi, et al.
Pubblicazione: (2025)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
di: Han, Wooseok, et al.
Pubblicazione: (2024)
di: Han, Wooseok, et al.
Pubblicazione: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
di: Ren, Yong, et al.
Pubblicazione: (2026)
di: Ren, Yong, et al.
Pubblicazione: (2026)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
Zero-Shot Mono-to-Binaural Speech Synthesis
di: Levkovitch, Alon, et al.
Pubblicazione: (2024)
di: Levkovitch, Alon, et al.
Pubblicazione: (2024)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
di: Cho, Chanhee, et al.
Pubblicazione: (2026)
di: Cho, Chanhee, et al.
Pubblicazione: (2026)
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
di: Han, Changjin, et al.
Pubblicazione: (2024)
di: Han, Changjin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2025) -
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024) -
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
di: Huang, Jiawei, et al.
Pubblicazione: (2024) -
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
di: Liu, Huadai, et al.
Pubblicazione: (2023) -
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
di: Zuo, Jialong, et al.
Pubblicazione: (2025)