Meta Learning Text-to-Speech Synthesis in over 7000 Languages
Fuente:
arXiv
Salvato in:
| Autori principali: | Lux, Florian, Meyer, Sarina, Behringer, Lyonel, Zalkow, Frank, Do, Phat, Coler, Matt, Habets, Emanuël A. P., Vu, Ngoc Thang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Probing the Feasibility of Multilingual Speaker Anonymization
di: Meyer, Sarina, et al.
Pubblicazione: (2024)
di: Meyer, Sarina, et al.
Pubblicazione: (2024)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
di: Bott, Thomas, et al.
Pubblicazione: (2024)
di: Bott, Thomas, et al.
Pubblicazione: (2024)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
di: Brendel, Andreas, et al.
Pubblicazione: (2024)
di: Brendel, Andreas, et al.
Pubblicazione: (2024)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
di: Denisov, Pavel, et al.
Pubblicazione: (2024)
di: Denisov, Pavel, et al.
Pubblicazione: (2024)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
High-Resolution Speech Restoration with Latent Diffusion Model
di: Dhyani, Tushar, et al.
Pubblicazione: (2024)
di: Dhyani, Tushar, et al.
Pubblicazione: (2024)
First Steps Towards Voice Anonymization for Code-Switching Speech
di: Meyer, Sarina, et al.
Pubblicazione: (2025)
di: Meyer, Sarina, et al.
Pubblicazione: (2025)
Use Cases for Voice Anonymization
di: Meyer, Sarina, et al.
Pubblicazione: (2025)
di: Meyer, Sarina, et al.
Pubblicazione: (2025)
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
di: Lakshminarayana, Kishor Kayyar, et al.
Pubblicazione: (2025)
di: Lakshminarayana, Kishor Kayyar, et al.
Pubblicazione: (2025)
Benchmarking Neural Speech Codec Intelligibility with SITool
di: Leschanowsky, Anna, et al.
Pubblicazione: (2025)
di: Leschanowsky, Anna, et al.
Pubblicazione: (2025)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
Stereo Reproduction in the Presence of Sample Rate Offsets
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
di: Li, Chia-Yu, et al.
Pubblicazione: (2024)
di: Li, Chia-Yu, et al.
Pubblicazione: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Data-driven Joint Detection and Localization of Acoustic Reflectors
di: Bicer, H. Nazim, et al.
Pubblicazione: (2024)
di: Bicer, H. Nazim, et al.
Pubblicazione: (2024)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2024)
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2024)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
di: Behringer, Lyonel, et al.
Pubblicazione: (2026)
di: Behringer, Lyonel, et al.
Pubblicazione: (2026)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
di: Wechsler, Julian, et al.
Pubblicazione: (2024)
di: Wechsler, Julian, et al.
Pubblicazione: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
di: Li, Zhu, et al.
Pubblicazione: (2024)
di: Li, Zhu, et al.
Pubblicazione: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
di: Liao, Shijia, et al.
Pubblicazione: (2024)
di: Liao, Shijia, et al.
Pubblicazione: (2024)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
di: Seki, Kentaro, et al.
Pubblicazione: (2025)
di: Seki, Kentaro, et al.
Pubblicazione: (2025)
A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
di: Cao, Junjie, et al.
Pubblicazione: (2025)
di: Cao, Junjie, et al.
Pubblicazione: (2025)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
di: Tu, Zehai, et al.
Pubblicazione: (2024)
di: Tu, Zehai, et al.
Pubblicazione: (2024)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
di: Wang, Helin, et al.
Pubblicazione: (2024)
di: Wang, Helin, et al.
Pubblicazione: (2024)
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model
di: Wang, Siyang, et al.
Pubblicazione: (2024)
di: Wang, Siyang, et al.
Pubblicazione: (2024)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
di: Ahmad, Hawraz A., et al.
Pubblicazione: (2024)
di: Ahmad, Hawraz A., et al.
Pubblicazione: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
di: Du, Chenpeng, et al.
Pubblicazione: (2022)
di: Du, Chenpeng, et al.
Pubblicazione: (2022)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Probing the Feasibility of Multilingual Speaker Anonymization
di: Meyer, Sarina, et al.
Pubblicazione: (2024) -
Controlling Emotion in Text-to-Speech with Natural Language Prompts
di: Bott, Thomas, et al.
Pubblicazione: (2024) -
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
di: Brendel, Andreas, et al.
Pubblicazione: (2024) -
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
di: Denisov, Pavel, et al.
Pubblicazione: (2024) -
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
di: Li, Zhu, et al.
Pubblicazione: (2025)