REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913979701395456 |
|---|---|
| author | Jiang, Yuepeng Ning, Ziqian Wang, Shuai Wang, Chengjia Bi, Mengxiao Zhu, Pengcheng Fu, Zhonghua Xie, Lei |
| author_facet | Jiang, Yuepeng Ning, Ziqian Wang, Shuai Wang, Chengjia Bi, Mengxiao Zhu, Pengcheng Fu, Zhonghua Xie, Lei |
| contents | In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL features, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that REF-VC outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_04996 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers Jiang, Yuepeng Ning, Ziqian Wang, Shuai Wang, Chengjia Bi, Mengxiao Zhu, Pengcheng Fu, Zhonghua Xie, Lei Audio and Speech Processing In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL features, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that REF-VC outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model. |
| title | REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.04996 |