Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Bae, Jae-Sung, Lee, Joun Yeop, Lee, Ji-Hyun, Mun, Seongkyu, Kang, Taehwa, Cho, Hoon-Young, Kim, Chanwoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
di: Lee, Joun Yeop, et al.
Pubblicazione: (2024)
di: Lee, Joun Yeop, et al.
Pubblicazione: (2024)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
di: Bae, Jae-Sung, et al.
Pubblicazione: (2025)
di: Bae, Jae-Sung, et al.
Pubblicazione: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
di: Kong, Jungil, et al.
Pubblicazione: (2023)
di: Kong, Jungil, et al.
Pubblicazione: (2023)
Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces
di: Coletta, Emma, et al.
Pubblicazione: (2026)
di: Coletta, Emma, et al.
Pubblicazione: (2026)
Efficient Parallel Audio Generation using Group Masked Language Modeling
di: Jeong, Myeonghun, et al.
Pubblicazione: (2024)
di: Jeong, Myeonghun, et al.
Pubblicazione: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
di: Li, Xiquan, et al.
Pubblicazione: (2024)
di: Li, Xiquan, et al.
Pubblicazione: (2024)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
di: Lee, Myungjin, et al.
Pubblicazione: (2026)
di: Lee, Myungjin, et al.
Pubblicazione: (2026)
VoxSim: A perceptual voice similarity dataset
di: Ahn, Junseok, et al.
Pubblicazione: (2024)
di: Ahn, Junseok, et al.
Pubblicazione: (2024)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
di: Lee, Yongjoon, et al.
Pubblicazione: (2024)
di: Lee, Yongjoon, et al.
Pubblicazione: (2024)
Single-Channel Distance-Based Source Separation for Mobile GPU in Outdoor and Indoor Environments
di: Bae, Hanbin, et al.
Pubblicazione: (2025)
di: Bae, Hanbin, et al.
Pubblicazione: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
di: Niu, Zhikang, et al.
Pubblicazione: (2025)
di: Niu, Zhikang, et al.
Pubblicazione: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2026)
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2026)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
di: Lin, Weiwei, et al.
Pubblicazione: (2024)
di: Lin, Weiwei, et al.
Pubblicazione: (2024)
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
di: Turetzky, Arnon, et al.
Pubblicazione: (2024)
di: Turetzky, Arnon, et al.
Pubblicazione: (2024)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
di: Zhu, Xinfa, et al.
Pubblicazione: (2025)
di: Zhu, Xinfa, et al.
Pubblicazione: (2025)
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
di: Wu, Chun Yat, et al.
Pubblicazione: (2025)
di: Wu, Chun Yat, et al.
Pubblicazione: (2025)
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
di: Huang, Wen, et al.
Pubblicazione: (2025)
di: Huang, Wen, et al.
Pubblicazione: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
di: Wang, Helin, et al.
Pubblicazione: (2024)
di: Wang, Helin, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2025)
di: Jiang, Ziyue, et al.
Pubblicazione: (2025)
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
di: Li, Jiatong, et al.
Pubblicazione: (2025)
di: Li, Jiatong, et al.
Pubblicazione: (2025)
AdaptVC: High Quality Voice Conversion with Adaptive Learning
di: Kim, Jaehun, et al.
Pubblicazione: (2025)
di: Kim, Jaehun, et al.
Pubblicazione: (2025)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
di: Xin, Detai, et al.
Pubblicazione: (2026)
di: Xin, Detai, et al.
Pubblicazione: (2026)
Naturalness-Aware Curriculum Learning with Dynamic Temperature for Speech Deepfake Detection
di: Kim, Taewoo, et al.
Pubblicazione: (2025)
di: Kim, Taewoo, et al.
Pubblicazione: (2025)
Instance-Specific Test-Time Training for Speech Editing in the Wild
di: Kim, Taewoo, et al.
Pubblicazione: (2025)
di: Kim, Taewoo, et al.
Pubblicazione: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2025)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2025)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2025)
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2025)
Latent CLAP Loss for Better Foley Sound Synthesis
di: Karchkhadze, Tornike, et al.
Pubblicazione: (2024)
di: Karchkhadze, Tornike, et al.
Pubblicazione: (2024)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
di: Yun, Minhyeok, et al.
Pubblicazione: (2026)
di: Yun, Minhyeok, et al.
Pubblicazione: (2026)
Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
di: Kim, Taewoo, et al.
Pubblicazione: (2024)
di: Kim, Taewoo, et al.
Pubblicazione: (2024)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
di: Kim, Tae-Woo, et al.
Pubblicazione: (2022)
di: Kim, Tae-Woo, et al.
Pubblicazione: (2022)
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
di: Han, Changjin, et al.
Pubblicazione: (2024)
di: Han, Changjin, et al.
Pubblicazione: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
di: Geng, Haopeng, et al.
Pubblicazione: (2024)
di: Geng, Haopeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
di: Lee, Joun Yeop, et al.
Pubblicazione: (2024) -
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
di: Kim, Minchan, et al.
Pubblicazione: (2024) -
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
di: Bae, Jae-Sung, et al.
Pubblicazione: (2025) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023) -
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
di: Kong, Jungil, et al.
Pubblicazione: (2023)