Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
Fuente:
arXiv
Saved in:
| Main Author: | Nippert, Lars |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
by: Răgman, Teodora, et al.
Published: (2024)
by: Răgman, Teodora, et al.
Published: (2024)
Incremental FastPitch: Chunk-based High Quality Text to Speech
by: Du, Muyang, et al.
Published: (2024)
by: Du, Muyang, et al.
Published: (2024)
E1 TTS: Simple and Fast Non-Autoregressive TTS
by: Liu, Zhijun, et al.
Published: (2024)
by: Liu, Zhijun, et al.
Published: (2024)
SwiftF0: Fast and Accurate Monophonic Pitch Detection
by: Nieradzik, Lars
Published: (2025)
by: Nieradzik, Lars
Published: (2025)
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
by: Zhao, Yuxiang, et al.
Published: (2025)
by: Zhao, Yuxiang, et al.
Published: (2025)
Periodicity Pitch Detection in Complex Harmonies on EEG Timeline Data
by: Heinze, Maria, et al.
Published: (2020)
by: Heinze, Maria, et al.
Published: (2020)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
by: Terashima, Ryo, et al.
Published: (2025)
by: Terashima, Ryo, et al.
Published: (2025)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
by: Kim, Tae-Woo, et al.
Published: (2022)
by: Kim, Tae-Woo, et al.
Published: (2022)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
Automotive Sound Quality for EVs: Psychoacoustic Metrics with Reproducible AI/ML Baselines
by: Goswami, Mandip
Published: (2025)
by: Goswami, Mandip
Published: (2025)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
by: Răgman, Teodora, et al.
Published: (2026)
by: Răgman, Teodora, et al.
Published: (2026)
Scalable Controllable Accented TTS
by: Xinyuan, Henry Li, et al.
Published: (2025)
by: Xinyuan, Henry Li, et al.
Published: (2025)
Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
by: Giraldo, Jose, et al.
Published: (2026)
by: Giraldo, Jose, et al.
Published: (2026)
SponTTS: modeling and transferring spontaneous style for TTS
by: Li, Hanzhao, et al.
Published: (2023)
by: Li, Hanzhao, et al.
Published: (2023)
T5Gemma-TTS Technical Report
by: Arata, Chihiro, et al.
Published: (2026)
by: Arata, Chihiro, et al.
Published: (2026)
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
by: Chung, Woo-Jin, et al.
Published: (2023)
by: Chung, Woo-Jin, et al.
Published: (2023)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
by: Guo, Yinlin, et al.
Published: (2024)
by: Guo, Yinlin, et al.
Published: (2024)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
by: Ma, Linhan, et al.
Published: (2024)
by: Ma, Linhan, et al.
Published: (2024)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
by: Rautenberg, Frederik, et al.
Published: (2026)
by: Rautenberg, Frederik, et al.
Published: (2026)
Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets
by: Di Bernardo, Matías, et al.
Published: (2025)
by: Di Bernardo, Matías, et al.
Published: (2025)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
by: Ko, Myeongjin, et al.
Published: (2023)
by: Ko, Myeongjin, et al.
Published: (2023)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
by: Eskimez, Sefik Emre, et al.
Published: (2024)
by: Eskimez, Sefik Emre, et al.
Published: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
by: Qharabagh, Mahta Fetrat, et al.
Published: (2024)
by: Qharabagh, Mahta Fetrat, et al.
Published: (2024)
Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track
by: Sun, Renhe, et al.
Published: (2026)
by: Sun, Renhe, et al.
Published: (2026)
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
by: Huo, Yanru, et al.
Published: (2025)
by: Huo, Yanru, et al.
Published: (2025)
Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
by: Park, Hyun Jin, et al.
Published: (2024)
by: Park, Hyun Jin, et al.
Published: (2024)
Improving Neural Pitch Estimation with SWIPE Kernels
by: Marttila, David, et al.
Published: (2025)
by: Marttila, David, et al.
Published: (2025)
Cross-domain Neural Pitch and Periodicity Estimation
by: Morrison, Max, et al.
Published: (2023)
by: Morrison, Max, et al.
Published: (2023)
Towards Flow-Matching-based TTS without Classifier-Free Guidance
by: Liang, Yuzhe, et al.
Published: (2025)
by: Liang, Yuzhe, et al.
Published: (2025)
Towards Prosodically Informed Mizo TTS without Explicit Tone Markings
by: Mohanta, Abhijit, et al.
Published: (2026)
by: Mohanta, Abhijit, et al.
Published: (2026)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
by: Kim, Semin, et al.
Published: (2026)
by: Kim, Semin, et al.
Published: (2026)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
by: Xue, Heyang, et al.
Published: (2025)
by: Xue, Heyang, et al.
Published: (2025)
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
by: Meng, Qingliang, et al.
Published: (2025)
by: Meng, Qingliang, et al.
Published: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
by: An, Keyu, et al.
Published: (2025)
by: An, Keyu, et al.
Published: (2025)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
by: Abilbekov, Adal, et al.
Published: (2024)
by: Abilbekov, Adal, et al.
Published: (2024)
Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
by: Ouyang, Zhicheng, et al.
Published: (2026)
by: Ouyang, Zhicheng, et al.
Published: (2026)
T-Mimi: A Transformer-based Mimi Decoder for Real-Time On-Phone TTS
by: Wu, Haibin, et al.
Published: (2026)
by: Wu, Haibin, et al.
Published: (2026)
Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
by: Valin, Jean-Marc, et al.
Published: (2024)
by: Valin, Jean-Marc, et al.
Published: (2024)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
by: Wang, Jianzong, et al.
Published: (2024)
by: Wang, Jianzong, et al.
Published: (2024)
Score-Based Training for Energy-Based TTS Models
by: Sun, Wanli, et al.
Published: (2025)
by: Sun, Wanli, et al.
Published: (2025)
Similar Items
-
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
by: Răgman, Teodora, et al.
Published: (2024) -
Incremental FastPitch: Chunk-based High Quality Text to Speech
by: Du, Muyang, et al.
Published: (2024) -
E1 TTS: Simple and Fast Non-Autoregressive TTS
by: Liu, Zhijun, et al.
Published: (2024) -
SwiftF0: Fast and Accurate Monophonic Pitch Detection
by: Nieradzik, Lars
Published: (2025) -
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
by: Zhao, Yuxiang, et al.
Published: (2025)