Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Plaja-Roglans, Genís, Hung, Yun-Ning, Serra, Xavier, Pereira, Igor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
von: Plaja-Roglans, Genís, et al.
Veröffentlicht: (2025)
von: Plaja-Roglans, Genís, et al.
Veröffentlicht: (2025)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
von: Sui, Kehan, et al.
Veröffentlicht: (2025)
von: Sui, Kehan, et al.
Veröffentlicht: (2025)
SingFake: Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis
von: Sui, Kehan, et al.
Veröffentlicht: (2024)
von: Sui, Kehan, et al.
Veröffentlicht: (2024)
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
von: Chen, Gongyu, et al.
Veröffentlicht: (2025)
von: Chen, Gongyu, et al.
Veröffentlicht: (2025)
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
von: Ryu, Yerin, et al.
Veröffentlicht: (2025)
von: Ryu, Yerin, et al.
Veröffentlicht: (2025)
Deepfake Detection of Singing Voices With Whisper Encodings
von: Sharma, Falguni, et al.
Veröffentlicht: (2025)
von: Sharma, Falguni, et al.
Veröffentlicht: (2025)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
von: Dhar, Sandipan, et al.
Veröffentlicht: (2025)
von: Dhar, Sandipan, et al.
Veröffentlicht: (2025)
Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
von: Sorrenti, Adam
Veröffentlicht: (2024)
von: Sorrenti, Adam
Veröffentlicht: (2024)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
CoMoSVC: Consistency Model-based Singing Voice Conversion
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
An Agent-Based Framework for Automated Higher-Voice Harmony Generation
von: Ganapathy, Nia D'Souza, et al.
Veröffentlicht: (2025)
von: Ganapathy, Nia D'Souza, et al.
Veröffentlicht: (2025)
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
von: Qian, Jiale, et al.
Veröffentlicht: (2026)
von: Qian, Jiale, et al.
Veröffentlicht: (2026)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Efficient Long-Sequence Diffusion Modeling for Symbolic Music Generation
von: Xu, Jinhan, et al.
Veröffentlicht: (2026)
von: Xu, Jinhan, et al.
Veröffentlicht: (2026)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
Towards Real-Time Human-AI Musical Co-Performance: Accompaniment Generation with Latent Diffusion Models and MAX/MSP
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2026)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2026)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
von: Choi, Joon-Seung, et al.
Veröffentlicht: (2025)
von: Choi, Joon-Seung, et al.
Veröffentlicht: (2025)
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
von: Hu, Zhetao, et al.
Veröffentlicht: (2026)
von: Hu, Zhetao, et al.
Veröffentlicht: (2026)
Latent-Mark: An Audio Watermark Robust to Neural Resynthesis
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
SLM-SS: Speech Language Model for Generative Speech Separation
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
von: Plaja-Roglans, Genís, et al.
Veröffentlicht: (2025) -
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026) -
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
von: Sui, Kehan, et al.
Veröffentlicht: (2025) -
SingFake: Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2023) -
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)