Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Shuqi, Wang, Yunyun, Dannenberg, Roger B., Jin, Zeyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909464618074112
author Dai, Shuqi
Wang, Yunyun
Dannenberg, Roger B.
Jin, Zeyu
author_facet Dai, Shuqi
Wang, Yunyun
Dannenberg, Roger B.
Jin, Zeyu
contents We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables control over multiple aspects, including language content based on lyrics, performance attributes based on a musical score, singing style and vocal techniques based on a selector, and voice identity based on a speech sample. The proposed zero-shot learning paradigm consists of one SVS model and two SVC models, utilizing pre-trained content embeddings and a diffusion-based generator. The proposed framework is also trained on mixed datasets comprising both singing and speech audio, allowing singing voice cloning based on speech reference. Experiments show substantial improvements in timbre similarity and musicality over state-of-the-art baselines, providing insights into other low-data music tasks such as instrumental style transfer. Examples can be found at: everyone-can-sing.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
Dai, Shuqi
Wang, Yunyun
Dannenberg, Roger B.
Jin, Zeyu
Sound
Audio and Speech Processing
We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables control over multiple aspects, including language content based on lyrics, performance attributes based on a musical score, singing style and vocal techniques based on a selector, and voice identity based on a speech sample. The proposed zero-shot learning paradigm consists of one SVS model and two SVC models, utilizing pre-trained content embeddings and a diffusion-based generator. The proposed framework is also trained on mixed datasets comprising both singing and speech audio, allowing singing voice cloning based on speech reference. Experiments show substantial improvements in timbre similarity and musicality over state-of-the-art baselines, providing insights into other low-data music tasks such as instrumental style transfer. Examples can be found at: everyone-can-sing.github.io.
title Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.13870