Accompanied Singing Voice Synthesis with Fully Text-controlled Melody

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Ruiqi, Hong, Zhiqing, Wang, Yongqi, Zhang, Lichao, Huang, Rongjie, Zheng, Siqi, Zhao, Zhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914855740506112
author Li, Ruiqi
Hong, Zhiqing
Wang, Yongqi
Zhang, Lichao
Huang, Rongjie
Zheng, Siqi
Zhao, Zhou
author_facet Li, Ruiqi
Hong, Zhiqing
Wang, Yongqi
Zhang, Lichao
Huang, Rongjie
Zheng, Siqi
Zhao, Zhou
contents Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be impractical, such as music scores or MIDI sequences. We present MelodyLM, the first TTSong model that generates high-quality song pieces with fully text-controlled melodies, achieving minimal user requirements and maximum control flexibility. MelodyLM explicitly models MIDI as the intermediate melody-related feature and sequentially generates vocal tracks in a language model manner, conditioned on textual and vocal prompts. The accompaniment music is subsequently synthesized by a latent diffusion model with hybrid conditioning for temporal alignment. With minimal requirements, users only need to input lyrics and a reference voice to synthesize a song sample. For full control, just input textual prompts or even directly input MIDI. Experimental results indicate that MelodyLM achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://melodylm666.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
Li, Ruiqi
Hong, Zhiqing
Wang, Yongqi
Zhang, Lichao
Huang, Rongjie
Zheng, Siqi
Zhao, Zhou
Audio and Speech Processing
Computation and Language
Sound
Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be impractical, such as music scores or MIDI sequences. We present MelodyLM, the first TTSong model that generates high-quality song pieces with fully text-controlled melodies, achieving minimal user requirements and maximum control flexibility. MelodyLM explicitly models MIDI as the intermediate melody-related feature and sequentially generates vocal tracks in a language model manner, conditioned on textual and vocal prompts. The accompaniment music is subsequently synthesized by a latent diffusion model with hybrid conditioning for temporal alignment. With minimal requirements, users only need to input lyrics and a reference voice to synthesize a song sample. For full control, just input textual prompts or even directly input MIDI. Experimental results indicate that MelodyLM achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://melodylm666.github.io.
title Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2407.02049