MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Semin, Jeong, Myeonghun, Lee, Hyeonseung, Kim, Minchan, Choi, Byoung Jin, Kim, Nam Soo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
Efficient Parallel Audio Generation using Group Masked Language Modeling
von: Jeong, Myeonghun, et al.
Veröffentlicht: (2024)
von: Jeong, Myeonghun, et al.
Veröffentlicht: (2024)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024)
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024)
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
von: Lee, Hyeonseung, et al.
Veröffentlicht: (2024)
von: Lee, Hyeonseung, et al.
Veröffentlicht: (2024)
SingIt! Singer Voice Transformation
von: Eliav, Amit, et al.
Veröffentlicht: (2024)
von: Eliav, Amit, et al.
Veröffentlicht: (2024)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
BiSinger: Bilingual Singing Voice Synthesis
von: Zhou, Huali, et al.
Veröffentlicht: (2023)
von: Zhou, Huali, et al.
Veröffentlicht: (2023)
PolySinger: Singing-Voice to Singing-Voice Translation from English to Japanese
von: Antonisen, Silas, et al.
Veröffentlicht: (2024)
von: Antonisen, Silas, et al.
Veröffentlicht: (2024)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
von: Pan, Changhao, et al.
Veröffentlicht: (2026)
von: Pan, Changhao, et al.
Veröffentlicht: (2026)
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
FxSearcher: gradient-free text-driven audio transformation
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
von: Qian, Jiale, et al.
Veröffentlicht: (2026)
von: Qian, Jiale, et al.
Veröffentlicht: (2026)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
von: Ryu, Yerin, et al.
Veröffentlicht: (2025)
von: Ryu, Yerin, et al.
Veröffentlicht: (2025)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
TokSing: Singing Voice Synthesis based on Discrete Tokens
von: Wu, Yuning, et al.
Veröffentlicht: (2024)
von: Wu, Yuning, et al.
Veröffentlicht: (2024)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
von: Song, Yulin, et al.
Veröffentlicht: (2024)
von: Song, Yulin, et al.
Veröffentlicht: (2024)
A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere In Between?
von: Choi, Kahyun, et al.
Veröffentlicht: (2024)
von: Choi, Kahyun, et al.
Veröffentlicht: (2024)
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
von: Kang, Ju Yeon, et al.
Veröffentlicht: (2025)
von: Kang, Ju Yeon, et al.
Veröffentlicht: (2025)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
Robust Singing Voice Transcription Serves Synthesis
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis
von: Sui, Kehan, et al.
Veröffentlicht: (2024)
von: Sui, Kehan, et al.
Veröffentlicht: (2024)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
von: Kim, Semin, et al.
Veröffentlicht: (2026)
von: Kim, Semin, et al.
Veröffentlicht: (2026)
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024) -
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
von: Kim, Minchan, et al.
Veröffentlicht: (2024) -
Efficient Parallel Audio Generation using Group Masked Language Modeling
von: Jeong, Myeonghun, et al.
Veröffentlicht: (2024) -
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024) -
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)