Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Zhijun, Wang, Shuai, Inoue, Sho, Bai, Qibing, Li, Haizhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
di: Jia, Dongya, et al.
Pubblicazione: (2025)
di: Jia, Dongya, et al.
Pubblicazione: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
Controllable Accent Normalization via Discrete Diffusion
di: Bai, Qibing, et al.
Pubblicazione: (2026)
di: Bai, Qibing, et al.
Pubblicazione: (2026)
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
di: Bai, Qibing, et al.
Pubblicazione: (2026)
di: Bai, Qibing, et al.
Pubblicazione: (2026)
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
di: Bai, Qibing, et al.
Pubblicazione: (2025)
di: Bai, Qibing, et al.
Pubblicazione: (2025)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
di: Ju, Zeqian, et al.
Pubblicazione: (2024)
di: Ju, Zeqian, et al.
Pubblicazione: (2024)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
di: Lee, Keon, et al.
Pubblicazione: (2024)
di: Lee, Keon, et al.
Pubblicazione: (2024)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
di: Toyin, Hawau Olamide, et al.
Pubblicazione: (2024)
di: Toyin, Hawau Olamide, et al.
Pubblicazione: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
di: Xin, Detai, et al.
Pubblicazione: (2024)
di: Xin, Detai, et al.
Pubblicazione: (2024)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
di: Lou, Haowei, et al.
Pubblicazione: (2024)
di: Lou, Haowei, et al.
Pubblicazione: (2024)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
di: Ye, Zhen, et al.
Pubblicazione: (2024)
di: Ye, Zhen, et al.
Pubblicazione: (2024)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
di: Zhou, Li, et al.
Pubblicazione: (2026)
di: Zhou, Li, et al.
Pubblicazione: (2026)
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
di: Battenberg, Eric, et al.
Pubblicazione: (2024)
di: Battenberg, Eric, et al.
Pubblicazione: (2024)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
di: Cai, Runyuan, et al.
Pubblicazione: (2026)
di: Cai, Runyuan, et al.
Pubblicazione: (2026)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
di: Li, Zehua Kcriss, et al.
Pubblicazione: (2024)
di: Li, Zehua Kcriss, et al.
Pubblicazione: (2024)
Scaling Speech Tokenizers with Diffusion Autoencoders
di: Wang, Yuancheng, et al.
Pubblicazione: (2026)
di: Wang, Yuancheng, et al.
Pubblicazione: (2026)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
di: Inoue, Nakamasa, et al.
Pubblicazione: (2024)
di: Inoue, Nakamasa, et al.
Pubblicazione: (2024)
Scaling Rich Style-Prompted Text-to-Speech Datasets
di: Diwan, Anuj, et al.
Pubblicazione: (2025)
di: Diwan, Anuj, et al.
Pubblicazione: (2025)
Collaborative Watermarking for Adversarial Speech Synthesis
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
di: Weise, Tobias, et al.
Pubblicazione: (2024)
di: Weise, Tobias, et al.
Pubblicazione: (2024)
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
di: Nam, KiHyun, et al.
Pubblicazione: (2025)
di: Nam, KiHyun, et al.
Pubblicazione: (2025)
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
di: Xie, Tianxin, et al.
Pubblicazione: (2024)
di: Xie, Tianxin, et al.
Pubblicazione: (2024)
Roadmap towards Superhuman Speech Understanding using Large Language Models
di: Bu, Fan, et al.
Pubblicazione: (2024)
di: Bu, Fan, et al.
Pubblicazione: (2024)
E1 TTS: Simple and Fast Non-Autoregressive TTS
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)
di: Lin, Zijian, et al.
Pubblicazione: (2025)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
di: Lou, Haowei, et al.
Pubblicazione: (2024)
di: Lou, Haowei, et al.
Pubblicazione: (2024)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
di: Wang, Wupeng, et al.
Pubblicazione: (2025)
di: Wang, Wupeng, et al.
Pubblicazione: (2025)
Masked Audio Generation using a Single Non-Autoregressive Transformer
di: Ziv, Alon, et al.
Pubblicazione: (2024)
di: Ziv, Alon, et al.
Pubblicazione: (2024)
Mitigating Unauthorized Speech Synthesis for Voice Protection
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
Decoding Order Matters in Autoregressive Speech Synthesis
di: Zhao, Minghui, et al.
Pubblicazione: (2026)
di: Zhao, Minghui, et al.
Pubblicazione: (2026)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024) -
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
di: Jia, Dongya, et al.
Pubblicazione: (2025) -
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025) -
Controllable Accent Normalization via Discrete Diffusion
di: Bai, Qibing, et al.
Pubblicazione: (2026) -
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
di: Bai, Qibing, et al.
Pubblicazione: (2026)