Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xueyao, Fang, Zihao, Gu, Yicheng, Chen, Haopeng, Zou, Lexiao, Zhang, Junan, Xue, Liumeng, Wu, Zhizheng
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917775931342848
author Zhang, Xueyao
Fang, Zihao
Gu, Yicheng
Chen, Haopeng
Zou, Lexiao
Zhang, Junan
Xue, Liumeng
Wu, Zhizheng
author_facet Zhang, Xueyao
Fang, Zihao
Gu, Yicheng
Chen, Haopeng
Zou, Lexiao
Zhang, Junan
Xue, Liumeng
Wu, Zhizheng
contents Singing Voice Conversion (SVC) is a technique that enables any singer to perform any song. To achieve this, it is essential to obtain speaker-agnostic representations from the source audio, which poses a significant challenge. A common solution involves utilizing a semantic-based audio pretrained model as a feature extractor. However, the degree to which the extracted features can meet the SVC requirements remains an open question. This includes their capability to accurately model melody and lyrics, the speaker-independency of their underlying acoustic information, and their robustness for in-the-wild acoustic environments. In this study, we investigate the knowledge within classical semantic-based pretrained models in much detail. We discover that the knowledge of different models is diverse and can be complementary for SVC. Based on the above, we design a Singing Voice Conversion framework based on Diverse Semantic-based Feature Fusion (DSFF-SVC). Experimental results demonstrate that DSFF-SVC can be generalized and improve various existing SVC models, particularly in challenging real-world conversion tasks. Our demo website is available at https://diversesemanticsvc.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2310_11160
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
Zhang, Xueyao
Fang, Zihao
Gu, Yicheng
Chen, Haopeng
Zou, Lexiao
Zhang, Junan
Xue, Liumeng
Wu, Zhizheng
Sound
Audio and Speech Processing
Singing Voice Conversion (SVC) is a technique that enables any singer to perform any song. To achieve this, it is essential to obtain speaker-agnostic representations from the source audio, which poses a significant challenge. A common solution involves utilizing a semantic-based audio pretrained model as a feature extractor. However, the degree to which the extracted features can meet the SVC requirements remains an open question. This includes their capability to accurately model melody and lyrics, the speaker-independency of their underlying acoustic information, and their robustness for in-the-wild acoustic environments. In this study, we investigate the knowledge within classical semantic-based pretrained models in much detail. We discover that the knowledge of different models is diverse and can be complementary for SVC. Based on the above, we design a Singing Voice Conversion framework based on Diverse Semantic-based Feature Fusion (DSFF-SVC). Experimental results demonstrate that DSFF-SVC can be generalized and improve various existing SVC models, particularly in challenging real-world conversion tasks. Our demo website is available at https://diversesemanticsvc.github.io/.
title Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2310.11160