VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Joon-Seung, Byun, Dong-Min, Oh, Hyung-Seok, Lee, Seong-Whan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914074407731200
author Choi, Joon-Seung
Byun, Dong-Min
Oh, Hyung-Seok
Lee, Seong-Whan
author_facet Choi, Joon-Seung
Byun, Dong-Min
Oh, Hyung-Seok
Lee, Seong-Whan
contents Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propose VibESVC, a controllable singing voice conversion model that explicitly extracts and manipulates vibrato using discrete wavelet transform. Unlike previous methods that model vibrato implicitly, our approach decomposes the F0 contour into frequency components, enabling precise transfer. This allows vibrato control for enhanced flexibility. Experimental results show that VibE-SVC effectively transforms singing styles while preserving speaker similarity. Both subjective and objective evaluations confirm high-quality conversion.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
Choi, Joon-Seung
Byun, Dong-Min
Oh, Hyung-Seok
Lee, Seong-Whan
Sound
Artificial Intelligence
Audio and Speech Processing
Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propose VibESVC, a controllable singing voice conversion model that explicitly extracts and manipulates vibrato using discrete wavelet transform. Unlike previous methods that model vibrato implicitly, our approach decomposes the F0 contour into frequency components, enabling precise transfer. This allows vibrato control for enhanced flexibility. Experimental results show that VibE-SVC effectively transforms singing styles while preserving speaker similarity. Both subjective and objective evaluations confirm high-quality conversion.
title VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.20794