Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Postolache, Emilian, Mariani, Giorgio, Cosmo, Luca, Benetos, Emmanouil, Rodolà, Emanuele
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914718541676544
author Postolache, Emilian
Mariani, Giorgio
Cosmo, Luca
Benetos, Emmanouil
Rodolà, Emanuele
author_facet Postolache, Emilian
Mariani, Giorgio
Cosmo, Luca
Benetos, Emmanouil
Rodolà, Emanuele
contents Multi-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separation. Despite their versatility, they require estimating the joint distribution over the sources, necessitating pre-separated musical data, which is rarely available, and fixing the number and type of sources at training time. This paper generalizes MSDM to arbitrary time-domain diffusion models conditioned on text embeddings. These models do not require separated data as they are trained on mixtures, can parameterize an arbitrary number of sources, and allow for rich semantic control. We propose an inference procedure enabling the coherent generation of sources and accompaniments. Additionally, we adapt the Dirac separator of MSDM to perform source separation. We experiment with diffusion models trained on Slakh2100 and MTG-Jamendo, showcasing competitive generation and separation results in a relaxed data setting.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models
Postolache, Emilian
Mariani, Giorgio
Cosmo, Luca
Benetos, Emmanouil
Rodolà, Emanuele
Sound
Machine Learning
Audio and Speech Processing
Multi-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separation. Despite their versatility, they require estimating the joint distribution over the sources, necessitating pre-separated musical data, which is rarely available, and fixing the number and type of sources at training time. This paper generalizes MSDM to arbitrary time-domain diffusion models conditioned on text embeddings. These models do not require separated data as they are trained on mixtures, can parameterize an arbitrary number of sources, and allow for rich semantic control. We propose an inference procedure enabling the coherent generation of sources and accompaniments. Additionally, we adapt the Dirac separator of MSDM to perform source separation. We experiment with diffusion models trained on Slakh2100 and MTG-Jamendo, showcasing competitive generation and separation results in a relaxed data setting.
title Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2403.11706