SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912640883752960 |
|---|---|
| author | Guo, Zhao Ning, Ziqian Ma, Guobin Xie, Lei |
| author_facet | Guo, Zhao Ning, Ziqian Ma, Guobin Xie, Lei |
| contents | Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_09245 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion Guo, Zhao Ning, Ziqian Ma, Guobin Xie, Lei Sound Audio and Speech Processing Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms. |
| title | SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2510.09245 |