StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yu, Huang, Rongjie, Li, Ruiqi, He, JinZheng, Xia, Yan, Chen, Feiyang, Duan, Xinyu, Huai, Baoxing, Zhao, Zhou
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912403930742784
author Zhang, Yu
Huang, Rongjie
Li, Ruiqi
He, JinZheng
Xia, Yan
Chen, Feiyang
Duan, Xinyu
Huai, Baoxing
Zhao, Zhou
author_facet Zhang, Yu
Huang, Rongjie
Li, Ruiqi
He, JinZheng
Xia, Yan
Chen, Feiyang
Duan, Xinyu
Huai, Baoxing
Zhao, Zhou
contents Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://aaronz345.github.io/StyleSingerDemo/.
format Preprint
id arxiv_https___arxiv_org_abs_2312_10741
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
Zhang, Yu
Huang, Rongjie
Li, Ruiqi
He, JinZheng
Xia, Yan
Chen, Feiyang
Duan, Xinyu
Huai, Baoxing
Zhao, Zhou
Audio and Speech Processing
Computation and Language
Sound
Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://aaronz345.github.io/StyleSingerDemo/.
title StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2312.10741