Saved in:
| Main Authors: | Plaja-Roglans, Genís, Hung, Yun-Ning, Serra, Xavier, Pereira, Igor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.21342 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
by: Plaja-Roglans, Genís, et al.
Published: (2025)
by: Plaja-Roglans, Genís, et al.
Published: (2025)
Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
by: Chen, Yukun, et al.
Published: (2026)
by: Chen, Yukun, et al.
Published: (2026)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
by: Watcharasupat, Karn N., et al.
Published: (2024)
by: Watcharasupat, Karn N., et al.
Published: (2024)
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
by: Sui, Kehan, et al.
Published: (2025)
by: Sui, Kehan, et al.
Published: (2025)
Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
by: Yun-Ning, et al.
Published: (2025)
by: Yun-Ning, et al.
Published: (2025)
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
by: Sorrenti, Adam
Published: (2024)
by: Sorrenti, Adam
Published: (2024)
YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
by: Chen, Gongyu, et al.
Published: (2025)
by: Chen, Gongyu, et al.
Published: (2025)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
by: Chen, Szu-Chi, et al.
Published: (2026)
by: Chen, Szu-Chi, et al.
Published: (2026)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
by: Batlle-Roca, Roser, et al.
Published: (2024)
by: Batlle-Roca, Roser, et al.
Published: (2024)
Hierarchical Generative Modeling of Melodic Vocal Contours in Hindustani Classical Music
by: Shikarpur, Nithya, et al.
Published: (2024)
by: Shikarpur, Nithya, et al.
Published: (2024)
Efficient Long-Sequence Diffusion Modeling for Symbolic Music Generation
by: Xu, Jinhan, et al.
Published: (2026)
by: Xu, Jinhan, et al.
Published: (2026)
YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
Sing it, Narrate it: Quality Musical Lyrics Translation
by: Ye, Zhuorui, et al.
Published: (2024)
by: Ye, Zhuorui, et al.
Published: (2024)
Music Style Transfer With Diffusion Model
by: Huang, Hong, et al.
Published: (2024)
by: Huang, Hong, et al.
Published: (2024)
A Conditioned UNet for Music Source Separation
by: O'Hanlon, Ken, et al.
Published: (2025)
by: O'Hanlon, Ken, et al.
Published: (2025)
HNote: Extending YNote with Hexadecimal Encoding for Fine-Tuning LLMs in Music Modeling
by: Chu, Hung-Ying, et al.
Published: (2025)
by: Chu, Hung-Ying, et al.
Published: (2025)
Device-Guided Music Transfer
by: Hung, Manh Pham, et al.
Published: (2025)
by: Hung, Manh Pham, et al.
Published: (2025)
MusicSynth: An Automated Pipeline for Generating Violin Fingerboard Animations from Sheet Music Using Optical Music Recognition
by: Kaushik, Abhimanyu
Published: (2026)
by: Kaushik, Abhimanyu
Published: (2026)
MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI
by: Batlle-Roca, Roser, et al.
Published: (2025)
by: Batlle-Roca, Roser, et al.
Published: (2025)
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
by: Novack, Zachary, et al.
Published: (2026)
by: Novack, Zachary, et al.
Published: (2026)
Robust Neural Audio Fingerprinting using Music Foundation Models
by: Singh, Shubhr, et al.
Published: (2025)
by: Singh, Shubhr, et al.
Published: (2025)
Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation
by: Cheng, Yuqing, et al.
Published: (2026)
by: Cheng, Yuqing, et al.
Published: (2026)
Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation
by: Ji, Shulei, et al.
Published: (2025)
by: Ji, Shulei, et al.
Published: (2025)
SingFake: Singing Voice Deepfake Detection
by: Zang, Yongyi, et al.
Published: (2023)
by: Zang, Yongyi, et al.
Published: (2023)
Composer Style-specific Symbolic Music Generation Using Vector Quantized Discrete Diffusion Models
by: Zhang, Jincheng, et al.
Published: (2023)
by: Zhang, Jincheng, et al.
Published: (2023)
Recognizing Ornaments in Vocal Indian Art Music with Active Annotation
by: Kumar, Sumit, et al.
Published: (2025)
by: Kumar, Sumit, et al.
Published: (2025)
Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation
by: Ebrahimzadeh, Maral, et al.
Published: (2025)
by: Ebrahimzadeh, Maral, et al.
Published: (2025)
Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
by: Kim, Euiyeon, et al.
Published: (2025)
by: Kim, Euiyeon, et al.
Published: (2025)
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
by: Morreale, Fabio, et al.
Published: (2025)
by: Morreale, Fabio, et al.
Published: (2025)
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis
by: Sui, Kehan, et al.
Published: (2024)
by: Sui, Kehan, et al.
Published: (2024)
Towards Real-Time Human-AI Musical Co-Performance: Accompaniment Generation with Latent Diffusion Models and MAX/MSP
by: Karchkhadze, Tornike, et al.
Published: (2026)
by: Karchkhadze, Tornike, et al.
Published: (2026)
Source Separation & Automatic Transcription for Music
by: Derby, Bradford, et al.
Published: (2024)
by: Derby, Bradford, et al.
Published: (2024)
SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
by: Tang, Yuxun, et al.
Published: (2025)
by: Tang, Yuxun, et al.
Published: (2025)
compIAM-ConvTDF-vocals-finetune
by: Schweinitz, Serafin, et al.
Published: (2025)
by: Schweinitz, Serafin, et al.
Published: (2025)
Mamba-Diffusion Model with Learnable Wavelet for Controllable Symbolic Music Generation
by: Zhang, Jincheng, et al.
Published: (2025)
by: Zhang, Jincheng, et al.
Published: (2025)
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
by: Hu, Zhetao, et al.
Published: (2026)
by: Hu, Zhetao, et al.
Published: (2026)
Training-Efficient Text-to-Music Generation with State-Space Modeling
by: Lee, Wei-Jaw, et al.
Published: (2026)
by: Lee, Wei-Jaw, et al.
Published: (2026)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
by: Yun, Minhyeok, et al.
Published: (2026)
by: Yun, Minhyeok, et al.
Published: (2026)
Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
by: Ramoneda, Pedro, et al.
Published: (2024)
by: Ramoneda, Pedro, et al.
Published: (2024)
Similar Items
-
Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
by: Plaja-Roglans, Genís, et al.
Published: (2025) -
Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition
by: Wang, Zihao, et al.
Published: (2025) -
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
by: Chen, Yukun, et al.
Published: (2026) -
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
by: Watcharasupat, Karn N., et al.
Published: (2024) -
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
by: Sui, Kehan, et al.
Published: (2025)