MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chae, Yunkee, Lee, Kyogu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912656349200384
author Chae, Yunkee
Lee, Kyogu
author_facet Chae, Yunkee
Lee, Kyogu
contents We present MGE-LDM, a unified latent diffusion framework for simultaneous music generation, source imputation, and query-driven source separation. Unlike prior approaches constrained to fixed instrument classes, MGE-LDM learns a joint distribution over full mixtures, submixtures, and individual stems within a single compact latent diffusion model. At inference, MGE-LDM enables (1) complete mixture generation, (2) partial generation (i.e., source imputation), and (3) text-conditioned extraction of arbitrary sources. By formulating both separation and imputation as conditional inpainting tasks in the latent space, our approach supports flexible, class-agnostic manipulation of arbitrary instrument sources. Notably, MGE-LDM can be trained jointly across heterogeneous multi-track datasets (e.g., Slakh2100, MUSDB18, MoisesDB) without relying on predefined instrument categories. Audio samples are available at our project page: https://yoongi43.github.io/MGELDM_Samples/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23305
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
Chae, Yunkee
Lee, Kyogu
Sound
Machine Learning
Audio and Speech Processing
We present MGE-LDM, a unified latent diffusion framework for simultaneous music generation, source imputation, and query-driven source separation. Unlike prior approaches constrained to fixed instrument classes, MGE-LDM learns a joint distribution over full mixtures, submixtures, and individual stems within a single compact latent diffusion model. At inference, MGE-LDM enables (1) complete mixture generation, (2) partial generation (i.e., source imputation), and (3) text-conditioned extraction of arbitrary sources. By formulating both separation and imputation as conditional inpainting tasks in the latent space, our approach supports flexible, class-agnostic manipulation of arbitrary instrument sources. Notably, MGE-LDM can be trained jointly across heterogeneous multi-track datasets (e.g., Slakh2100, MUSDB18, MoisesDB) without relying on predefined instrument categories. Audio samples are available at our project page: https://yoongi43.github.io/MGELDM_Samples/.
title MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2505.23305