Saved in:
| Main Authors: | Ai, Jiabao, Zhao, Minghui, Ragni, Anton |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.14032 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discrete-Time Diffusion-Like Models for Speech Synthesis
by: Tan, Xiaozhou, et al.
Published: (2025)
by: Tan, Xiaozhou, et al.
Published: (2025)
Decoding Order Matters in Autoregressive Speech Synthesis
by: Zhao, Minghui, et al.
Published: (2026)
by: Zhao, Minghui, et al.
Published: (2026)
Score-Based Training for Energy-Based TTS Models
by: Sun, Wanli, et al.
Published: (2025)
by: Sun, Wanli, et al.
Published: (2025)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
by: Flynn, Robert, et al.
Published: (2026)
by: Flynn, Robert, et al.
Published: (2026)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
by: Liu, Huadai, et al.
Published: (2023)
by: Liu, Huadai, et al.
Published: (2023)
How Much Context Does My Attention-Based ASR System Need?
by: Flynn, Robert, et al.
Published: (2023)
by: Flynn, Robert, et al.
Published: (2023)
Emphasis Sensitivity in Speech Representations
by: Cassini, Shaun, et al.
Published: (2025)
by: Cassini, Shaun, et al.
Published: (2025)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
by: He, Xinlu, et al.
Published: (2025)
by: He, Xinlu, et al.
Published: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
by: An, Keyu, et al.
Published: (2025)
by: An, Keyu, et al.
Published: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
by: Leung, Wing-Zin, et al.
Published: (2024)
by: Leung, Wing-Zin, et al.
Published: (2024)
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
by: Zhao, Yuxiang, et al.
Published: (2025)
by: Zhao, Yuxiang, et al.
Published: (2025)
Self-Train Before You Transcribe
by: Flynn, Robert, et al.
Published: (2024)
by: Flynn, Robert, et al.
Published: (2024)
Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation
by: Xu, Zhongweiyang, et al.
Published: (2026)
by: Xu, Zhongweiyang, et al.
Published: (2026)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
by: Lu, Ye-Xin, et al.
Published: (2025)
by: Lu, Ye-Xin, et al.
Published: (2025)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction
by: Huang, Weilong, et al.
Published: (2026)
by: Huang, Weilong, et al.
Published: (2026)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
by: Răgman, Teodora, et al.
Published: (2026)
by: Răgman, Teodora, et al.
Published: (2026)
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
by: Park, Hyun Joon, et al.
Published: (2024)
by: Park, Hyun Joon, et al.
Published: (2024)
Diffusion-based Signal Refiner for Speech Enhancement and Separation
by: Hirano, Masato, et al.
Published: (2023)
by: Hirano, Masato, et al.
Published: (2023)
Improving Music Source Separation with Diffusion and Consistency Refinement
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
by: Ma, Linhan, et al.
Published: (2024)
by: Ma, Linhan, et al.
Published: (2024)
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
by: Yang, Chenyu, et al.
Published: (2025)
by: Yang, Chenyu, et al.
Published: (2025)
SponTTS: modeling and transferring spontaneous style for TTS
by: Li, Hanzhao, et al.
Published: (2023)
by: Li, Hanzhao, et al.
Published: (2023)
Scalable Controllable Accented TTS
by: Xinyuan, Henry Li, et al.
Published: (2025)
by: Xinyuan, Henry Li, et al.
Published: (2025)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
by: Nie, Sihang, et al.
Published: (2025)
by: Nie, Sihang, et al.
Published: (2025)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
by: Chen, Zhiyong, et al.
Published: (2024)
by: Chen, Zhiyong, et al.
Published: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
by: Eskimez, Sefik Emre, et al.
Published: (2024)
by: Eskimez, Sefik Emre, et al.
Published: (2024)
E1 TTS: Simple and Fast Non-Autoregressive TTS
by: Liu, Zhijun, et al.
Published: (2024)
by: Liu, Zhijun, et al.
Published: (2024)
Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
by: Giraldo, Jose, et al.
Published: (2026)
by: Giraldo, Jose, et al.
Published: (2026)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
by: Lu, Ye-Xin, et al.
Published: (2025)
by: Lu, Ye-Xin, et al.
Published: (2025)
Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
by: Liu, Chenlin, et al.
Published: (2025)
by: Liu, Chenlin, et al.
Published: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
by: Xue, Heyang, et al.
Published: (2025)
by: Xue, Heyang, et al.
Published: (2025)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
by: Nguyen, Tan Dat, et al.
Published: (2025)
by: Nguyen, Tan Dat, et al.
Published: (2025)
T5Gemma-TTS Technical Report
by: Arata, Chihiro, et al.
Published: (2026)
by: Arata, Chihiro, et al.
Published: (2026)
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS
by: Seo, Deokjin, et al.
Published: (2026)
by: Seo, Deokjin, et al.
Published: (2026)
A Non-autoregressive Model for Joint STT and TTS
by: Sunder, Vishal, et al.
Published: (2025)
by: Sunder, Vishal, et al.
Published: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
by: Li, Yinghao Aaron, et al.
Published: (2024)
by: Li, Yinghao Aaron, et al.
Published: (2024)
ProSE: Diffusion Priors for Speech Enhancement
by: Kumar, Sonal, et al.
Published: (2025)
by: Kumar, Sonal, et al.
Published: (2025)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
by: You, Fuming, et al.
Published: (2024)
by: You, Fuming, et al.
Published: (2024)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2025)
by: Jiang, Ziyue, et al.
Published: (2025)
Similar Items
-
Discrete-Time Diffusion-Like Models for Speech Synthesis
by: Tan, Xiaozhou, et al.
Published: (2025) -
Decoding Order Matters in Autoregressive Speech Synthesis
by: Zhao, Minghui, et al.
Published: (2026) -
Score-Based Training for Energy-Based TTS Models
by: Sun, Wanli, et al.
Published: (2025) -
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
by: Flynn, Robert, et al.
Published: (2026) -
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
by: Liu, Huadai, et al.
Published: (2023)