Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Zhiqing, Huang, Rongjie, Cheng, Xize, Wang, Yongqi, Li, Ruiqi, You, Fuming, Zhao, Zhou, Zhang, Zhimeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913355750440960
author Hong, Zhiqing
Huang, Rongjie
Cheng, Xize
Wang, Yongqi
Li, Ruiqi
You, Fuming
Zhao, Zhou
Zhang, Zhimeng
author_facet Hong, Zhiqing
Huang, Rongjie
Cheng, Xize
Wang, Yongqi
Li, Ruiqi
You, Fuming
Zhao, Zhou
Zhang, Zhimeng
contents A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel task called text-to-song synthesis which incorporating both vocals and accompaniments generation. We develop Melodist, a two-stage text-to-song method that consists of singing voice synthesis (SVS) and vocal-to-accompaniment (V2A) synthesis. Melodist leverages tri-tower contrastive pretraining to learn more effective text representation for controllable V2A synthesis. A Chinese song dataset mined from a music website is built up to alleviate data scarcity for our research. The evaluation results on our dataset demonstrate that Melodist can synthesize songs with comparable quality and style consistency. Audio samples can be found in https://text2songMelodist.github.io/Sample/.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
Hong, Zhiqing
Huang, Rongjie
Cheng, Xize
Wang, Yongqi
Li, Ruiqi
You, Fuming
Zhao, Zhou
Zhang, Zhimeng
Audio and Speech Processing
Artificial Intelligence
A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel task called text-to-song synthesis which incorporating both vocals and accompaniments generation. We develop Melodist, a two-stage text-to-song method that consists of singing voice synthesis (SVS) and vocal-to-accompaniment (V2A) synthesis. Melodist leverages tri-tower contrastive pretraining to learn more effective text representation for controllable V2A synthesis. A Chinese song dataset mined from a music website is built up to alleviate data scarcity for our research. The evaluation results on our dataset demonstrate that Melodist can synthesize songs with comparable quality and style consistency. Audio samples can be found in https://text2songMelodist.github.io/Sample/.
title Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2404.09313