Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
Fuente:
arXiv
Guardado en:
| Autores principales: | Lakshminarayana, Kishor Kayyar, Zalkow, Frank, Dittmar, Christian, Pia, Nicola, Habets, Emanuel A. P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking Neural Speech Codec Intelligibility with SITool
por: Leschanowsky, Anna, et al.
Publicado: (2025)
por: Leschanowsky, Anna, et al.
Publicado: (2025)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
por: Lux, Florian, et al.
Publicado: (2024)
por: Lux, Florian, et al.
Publicado: (2024)
Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
Dynamic Slimmable Networks for Efficient Speech Separation
por: Elminshawi, Mohamed, et al.
Publicado: (2025)
por: Elminshawi, Mohamed, et al.
Publicado: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
por: Grundhuber, Philipp, et al.
Publicado: (2025)
por: Grundhuber, Philipp, et al.
Publicado: (2025)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
por: Halimeh, Mhd Modar, et al.
Publicado: (2025)
por: Halimeh, Mhd Modar, et al.
Publicado: (2025)
A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
por: Nespoli, Francesco, et al.
Publicado: (2024)
por: Nespoli, Francesco, et al.
Publicado: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
por: Leung, Wing-Zin, et al.
Publicado: (2024)
por: Leung, Wing-Zin, et al.
Publicado: (2024)
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
por: Korse, Srikanth, et al.
Publicado: (2025)
por: Korse, Srikanth, et al.
Publicado: (2025)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
Low-Complexity Neural Wind Noise Reduction for Audio Recordings
por: Eftekhari, Hesam, et al.
Publicado: (2025)
por: Eftekhari, Hesam, et al.
Publicado: (2025)
Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
por: Götz, Philipp, et al.
Publicado: (2025)
por: Götz, Philipp, et al.
Publicado: (2025)
Sample Rate Offset Compensated Acoustic Echo Cancellation For Multi-Device Scenarios
por: Korse, Srikanth, et al.
Publicado: (2025)
por: Korse, Srikanth, et al.
Publicado: (2025)
Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections
por: Xu, Zeyu, et al.
Publicado: (2026)
por: Xu, Zeyu, et al.
Publicado: (2026)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
por: Kuhlmann, Michael, et al.
Publicado: (2026)
por: Kuhlmann, Michael, et al.
Publicado: (2026)
Blind Acoustic Parameter Estimation Through Task-Agnostic Embeddings Using Latent Approximations
por: Götz, Philipp, et al.
Publicado: (2024)
por: Götz, Philipp, et al.
Publicado: (2024)
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
por: Halimeh, Mhd Modar, et al.
Publicado: (2024)
por: Halimeh, Mhd Modar, et al.
Publicado: (2024)
Neural Directional Filtering with Configurable Directivity Pattern at Inference
por: Huang, Weilong, et al.
Publicado: (2025)
por: Huang, Weilong, et al.
Publicado: (2025)
Navigating PESQ: Up-to-Date Versions and Open Implementations
por: Torcoli, Matteo, et al.
Publicado: (2025)
por: Torcoli, Matteo, et al.
Publicado: (2025)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
por: Grundhuber, Philipp, et al.
Publicado: (2025)
por: Grundhuber, Philipp, et al.
Publicado: (2025)
Speech Loudness in Broadcasting and Streaming
por: Torcoli, Matteo, et al.
Publicado: (2024)
por: Torcoli, Matteo, et al.
Publicado: (2024)
Stereo Reproduction in the Presence of Sample Rate Offsets
por: Korse, Srikanth, et al.
Publicado: (2025)
por: Korse, Srikanth, et al.
Publicado: (2025)
Neural Directional Filtering Using a Compact Microphone Array
por: Huang, Weilong, et al.
Publicado: (2025)
por: Huang, Weilong, et al.
Publicado: (2025)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
por: Aloradi, Ahmad, et al.
Publicado: (2025)
por: Aloradi, Ahmad, et al.
Publicado: (2025)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
por: Liang, Ziqi, et al.
Publicado: (2024)
por: Liang, Ziqi, et al.
Publicado: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
por: Yang, Yifan, et al.
Publicado: (2024)
por: Yang, Yifan, et al.
Publicado: (2024)
NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction
por: Huang, Weilong, et al.
Publicado: (2026)
por: Huang, Weilong, et al.
Publicado: (2026)
Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
por: Susladkar, Onkar Kishor, et al.
Publicado: (2024)
por: Susladkar, Onkar Kishor, et al.
Publicado: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
Data-driven Joint Detection and Localization of Acoustic Reflectors
por: Bicer, H. Nazim, et al.
Publicado: (2024)
por: Bicer, H. Nazim, et al.
Publicado: (2024)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
por: Gállego, Gerard I., et al.
Publicado: (2025)
por: Gállego, Gerard I., et al.
Publicado: (2025)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
por: Dao, Alan, et al.
Publicado: (2025)
por: Dao, Alan, et al.
Publicado: (2025)
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
por: Huang, Yiqiao, et al.
Publicado: (2024)
por: Huang, Yiqiao, et al.
Publicado: (2024)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
por: Chung, Raymond
Publicado: (2026)
por: Chung, Raymond
Publicado: (2026)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
por: Kim, Youngjae, et al.
Publicado: (2024)
por: Kim, Youngjae, et al.
Publicado: (2024)
UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension
por: Gupta, Kishan, et al.
Publicado: (2025)
por: Gupta, Kishan, et al.
Publicado: (2025)
Ejemplares similares
-
Benchmarking Neural Speech Codec Intelligibility with SITool
por: Leschanowsky, Anna, et al.
Publicado: (2025) -
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
por: Lux, Florian, et al.
Publicado: (2024) -
Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
por: Shetu, Shrishti Saha, et al.
Publicado: (2024) -
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
por: Shetu, Shrishti Saha, et al.
Publicado: (2024) -
Dynamic Slimmable Networks for Efficient Speech Separation
por: Elminshawi, Mohamed, et al.
Publicado: (2025)