R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Junjie, Chen, Gongyu, Ding, Chaofan, Chen, Zihao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914109416538112
author Zheng, Junjie
Chen, Gongyu
Ding, Chaofan
Chen, Zihao
author_facet Zheng, Junjie
Chen, Gongyu
Ding, Chaofan
Chen, Zihao
contents In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment scenarios, as both training and inference usually rely on clean data. This mismatch hinders practical use, given the inevitable presence of diverse noise sources and artifacts from music separation. To tackle these issues, we propose R2-SVC, a robust and expressive SVC framework. First, we introduce simulation-based robustness enhancement through random fundamental frequency ($F_0$) perturbations and music separation artifact simulations (e.g., reverberation, echo), substantially improving performance under noisy conditions. Second, we enrich speaker representation using domain-specific singing data: alongside clean vocals, we incorporate DNSMOS-filtered separated vocals and public singing corpora, enabling the model to preserve speaker timbre while capturing singing style nuances. Third, we integrate the Neural Source-Filter (NSF) model to explicitly represent harmonic and noise components, enhancing the naturalness and controllability of converted singing. R2-SVC achieves state-of-the-art results on multiple SVC benchmarks under both clean and noisy conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20677
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
Zheng, Junjie
Chen, Gongyu
Ding, Chaofan
Chen, Zihao
Sound
Artificial Intelligence
Audio and Speech Processing
In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment scenarios, as both training and inference usually rely on clean data. This mismatch hinders practical use, given the inevitable presence of diverse noise sources and artifacts from music separation. To tackle these issues, we propose R2-SVC, a robust and expressive SVC framework. First, we introduce simulation-based robustness enhancement through random fundamental frequency ($F_0$) perturbations and music separation artifact simulations (e.g., reverberation, echo), substantially improving performance under noisy conditions. Second, we enrich speaker representation using domain-specific singing data: alongside clean vocals, we incorporate DNSMOS-filtered separated vocals and public singing corpora, enabling the model to preserve speaker timbre while capturing singing style nuances. Third, we integrate the Neural Source-Filter (NSF) model to explicitly represent harmonic and noise components, enhancing the naturalness and controllability of converted singing. R2-SVC achieves state-of-the-art results on multiple SVC benchmarks under both clean and noisy conditions.
title R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.20677