SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909686365683712 |
|---|---|
| author | Mehta, Shivam Liu, Yingru Tang, Zhenyu Peng, Kainan Manohar, Vimal Zhang, Shun Seltzer, Mike He, Qing Ma, Mingbo |
| author_facet | Mehta, Shivam Liu, Yingru Tang, Zhenyu Peng, Kainan Manohar, Vimal Zhang, Shun Seltzer, Mike He, Qing Ma, Mingbo |
| contents | Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_09070 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment Mehta, Shivam Liu, Yingru Tang, Zhenyu Peng, Kainan Manohar, Vimal Zhang, Shun Seltzer, Mike He, Qing Ma, Mingbo Audio and Speech Processing Sound 68T07 I.2.7; I.2.6; G.3; H.5.5 Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/ |
| title | SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment |
| topic | Audio and Speech Processing Sound 68T07 I.2.7; I.2.6; G.3; H.5.5 |
| url | https://arxiv.org/abs/2507.09070 |