Melody-Guided Music Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wei, Shaopeng, Wei, Manzhen, Wang, Haoyu, Zhao, Yu, Kou, Gang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915083082268672
author Wei, Shaopeng
Wei, Manzhen
Wang, Haoyu
Zhao, Yu
Kou, Gang
author_facet Wei, Shaopeng
Wei, Manzhen
Wang, Haoyu
Zhao, Yu
Kou, Gang
contents We present the Melody-Guided Music Generation (MG2) model, a novel approach using melody to guide the text-to-music generation that, despite a simple method and limited resources, achieves excellent performance. Specifically, we first align the text with audio waveforms and their associated melodies using the newly proposed Contrastive Language-Music Pretraining, enabling the learned text representation fused with implicit melody information. Subsequently, we condition the retrieval-augmented diffusion module on both text prompt and retrieved melody. This allows MG2 to generate music that reflects the content of the given text description, meantime keeping the intrinsic harmony under the guidance of explicit melody information. We conducted extensive experiments on two public datasets: MusicCaps and MusicBench. Surprisingly, the experimental results demonstrate that the proposed MG2 model surpasses current open-source text-to-music generation models, achieving this with fewer than 1/3 of the parameters or less than 1/200 of the training data compared to state-of-the-art counterparts. Furthermore, we conducted comprehensive human evaluations involving three types of users and five perspectives, using newly designed questionnaires to explore the potential real-world applications of MG2.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20196
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Melody-Guided Music Generation
Wei, Shaopeng
Wei, Manzhen
Wang, Haoyu
Zhao, Yu
Kou, Gang
Sound
Artificial Intelligence
Audio and Speech Processing
We present the Melody-Guided Music Generation (MG2) model, a novel approach using melody to guide the text-to-music generation that, despite a simple method and limited resources, achieves excellent performance. Specifically, we first align the text with audio waveforms and their associated melodies using the newly proposed Contrastive Language-Music Pretraining, enabling the learned text representation fused with implicit melody information. Subsequently, we condition the retrieval-augmented diffusion module on both text prompt and retrieved melody. This allows MG2 to generate music that reflects the content of the given text description, meantime keeping the intrinsic harmony under the guidance of explicit melody information. We conducted extensive experiments on two public datasets: MusicCaps and MusicBench. Surprisingly, the experimental results demonstrate that the proposed MG2 model surpasses current open-source text-to-music generation models, achieving this with fewer than 1/3 of the parameters or less than 1/200 of the training data compared to state-of-the-art counterparts. Furthermore, we conducted comprehensive human evaluations involving three types of users and five perspectives, using newly designed questionnaires to explore the potential real-world applications of MG2.
title Melody-Guided Music Generation
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2409.20196