MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Ziyue, Ren, Yi, Li, Ruiqi, Ji, Shengpeng, Zhang, Boyang, Ye, Zhenhui, Zhang, Chen, Jionghao, Bai, Yang, Xiaoda, Zuo, Jialong, Zhang, Yu, Liu, Rui, Yin, Xiang, Zhao, Zhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
Intelli-Z: Toward Intelligible Zero-Shot TTS
di: Jung, Sunghee, et al.
Pubblicazione: (2024)
di: Jung, Sunghee, et al.
Pubblicazione: (2024)
Speech Watermarking with Discrete Intermediate Representations
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
di: Wang, Yongqi, et al.
Pubblicazione: (2023)
di: Wang, Yongqi, et al.
Pubblicazione: (2023)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
di: Ren, Yong, et al.
Pubblicazione: (2026)
di: Ren, Yong, et al.
Pubblicazione: (2026)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
di: Lam, Perry, et al.
Pubblicazione: (2022)
di: Lam, Perry, et al.
Pubblicazione: (2022)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
di: Li, Haitao, et al.
Pubblicazione: (2026)
di: Li, Haitao, et al.
Pubblicazione: (2026)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
di: Deng, Wei, et al.
Pubblicazione: (2025)
di: Deng, Wei, et al.
Pubblicazione: (2025)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
di: Chen, Ziqi, et al.
Pubblicazione: (2025)
di: Chen, Ziqi, et al.
Pubblicazione: (2025)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
di: Pankov, Vikentii, et al.
Pubblicazione: (2023)
di: Pankov, Vikentii, et al.
Pubblicazione: (2023)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
di: Arora, Akshit, et al.
Pubblicazione: (2024)
di: Arora, Akshit, et al.
Pubblicazione: (2024)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM
di: Yu, Jiawei, et al.
Pubblicazione: (2024)
di: Yu, Jiawei, et al.
Pubblicazione: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2023) -
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024) -
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
di: Zuo, Jialong, et al.
Pubblicazione: (2025) -
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024) -
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)