Gespeichert in:
| Hauptverfasser: | Giraldo, Jose, Peiró-Lilja, Alex, Zevallos, Rodolfo, España-Bonet, Cristina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.05770 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track
von: Sun, Renhe, et al.
Veröffentlicht: (2026)
von: Sun, Renhe, et al.
Veröffentlicht: (2026)
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024)
von: Giraldo, José, et al.
Veröffentlicht: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2025)
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Voice Impression Control in Zero-Shot TTS
von: Fujita, Kenichi, et al.
Veröffentlicht: (2025)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2025)
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
von: Ouyang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ouyang, Zhicheng, et al.
Veröffentlicht: (2026)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
von: Wang, Xiaofei, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofei, et al.
Veröffentlicht: (2024)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
von: Li, Zirui, et al.
Veröffentlicht: (2025)
von: Li, Zirui, et al.
Veröffentlicht: (2025)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
von: Zhou, Shuoyi, et al.
Veröffentlicht: (2024)
von: Zhou, Shuoyi, et al.
Veröffentlicht: (2024)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
von: Răgman, Teodora, et al.
Veröffentlicht: (2026)
von: Răgman, Teodora, et al.
Veröffentlicht: (2026)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
von: Manku, Ruskin Raj, et al.
Veröffentlicht: (2025)
von: Manku, Ruskin Raj, et al.
Veröffentlicht: (2025)
Scalable Controllable Accented TTS
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2025)
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2025)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS
von: Seo, Deokjin, et al.
Veröffentlicht: (2026)
von: Seo, Deokjin, et al.
Veröffentlicht: (2026)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2025)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2025)
T5Gemma-TTS Technical Report
von: Arata, Chihiro, et al.
Veröffentlicht: (2026)
von: Arata, Chihiro, et al.
Veröffentlicht: (2026)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
von: Pankov, Vikentii, et al.
Veröffentlicht: (2023)
von: Pankov, Vikentii, et al.
Veröffentlicht: (2023)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
von: Deng, Wei, et al.
Veröffentlicht: (2025)
von: Deng, Wei, et al.
Veröffentlicht: (2025)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track
von: Sun, Renhe, et al.
Veröffentlicht: (2026) -
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024) -
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024) -
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)