SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mehta, Shivam, Liu, Yingru, Tang, Zhenyu, Peng, Kainan, Manohar, Vimal, Zhang, Shun, Seltzer, Mike, He, Qing, Ma, Mingbo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909686365683712
author Mehta, Shivam
Liu, Yingru
Tang, Zhenyu
Peng, Kainan
Manohar, Vimal
Zhang, Shun
Seltzer, Mike
He, Qing
Ma, Mingbo
author_facet Mehta, Shivam
Liu, Yingru
Tang, Zhenyu
Peng, Kainan
Manohar, Vimal
Zhang, Shun
Seltzer, Mike
He, Qing
Ma, Mingbo
contents Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/
format Preprint
id arxiv_https___arxiv_org_abs_2507_09070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
Mehta, Shivam
Liu, Yingru
Tang, Zhenyu
Peng, Kainan
Manohar, Vimal
Zhang, Shun
Seltzer, Mike
He, Qing
Ma, Mingbo
Audio and Speech Processing
Sound
68T07
I.2.7; I.2.6; G.3; H.5.5
Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/
title SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
topic Audio and Speech Processing
Sound
68T07
I.2.7; I.2.6; G.3; H.5.5
url https://arxiv.org/abs/2507.09070