REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jiang, Yuepeng, Ning, Ziqian, Wang, Shuai, Wang, Chengjia, Bi, Mengxiao, Zhu, Pengcheng, Fu, Zhonghua, Xie, Lei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913979701395456
author Jiang, Yuepeng
Ning, Ziqian
Wang, Shuai
Wang, Chengjia
Bi, Mengxiao
Zhu, Pengcheng
Fu, Zhonghua
Xie, Lei
author_facet Jiang, Yuepeng
Ning, Ziqian
Wang, Shuai
Wang, Chengjia
Bi, Mengxiao
Zhu, Pengcheng
Fu, Zhonghua
Xie, Lei
contents In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL features, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that REF-VC outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
Jiang, Yuepeng
Ning, Ziqian
Wang, Shuai
Wang, Chengjia
Bi, Mengxiao
Zhu, Pengcheng
Fu, Zhonghua
Xie, Lei
Audio and Speech Processing
In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL features, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that REF-VC outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model.
title REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.04996