SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guo, Zhao, Ning, Ziqian, Ma, Guobin, Xie, Lei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912640883752960
author Guo, Zhao
Ning, Ziqian
Ma, Guobin
Xie, Lei
author_facet Guo, Zhao
Ning, Ziqian
Ma, Guobin
Xie, Lei
contents Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
Guo, Zhao
Ning, Ziqian
Ma, Guobin
Xie, Lei
Sound
Audio and Speech Processing
Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms.
title SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.09245