Training-Free Voice Conversion with Factorized Optimal Transport

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lobashev, Alexander, Yermekova, Assel, Larchenko, Maria
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916790650535936
author Lobashev, Alexander
Yermekova, Assel
Larchenko, Maria
author_facet Lobashev, Alexander
Yermekova, Assel
Larchenko, Maria
contents This paper introduces Factorized MKL-VC, a training-free modification for kNN-VC pipeline. In contrast with original pipeline, our algorithm performs high quality any-to-any cross-lingual voice conversion with only 5 second of reference audio. MKL-VC replaces kNN regression with a factorized optimal transport map in WavLM embedding subspaces, derived from Monge-Kantorovich Linear solution. Factorization addresses non-uniform variance across dimensions, ensuring effective feature transformation. Experiments on LibriSpeech and FLEURS datasets show MKL-VC significantly improves content preservation and robustness with short reference audio, outperforming kNN-VC. MKL-VC achieves performance comparable to FACodec, especially in cross-lingual voice conversion domain.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09709
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-Free Voice Conversion with Factorized Optimal Transport
Lobashev, Alexander
Yermekova, Assel
Larchenko, Maria
Sound
Computer Vision and Pattern Recognition
Machine Learning
Audio and Speech Processing
This paper introduces Factorized MKL-VC, a training-free modification for kNN-VC pipeline. In contrast with original pipeline, our algorithm performs high quality any-to-any cross-lingual voice conversion with only 5 second of reference audio. MKL-VC replaces kNN regression with a factorized optimal transport map in WavLM embedding subspaces, derived from Monge-Kantorovich Linear solution. Factorization addresses non-uniform variance across dimensions, ensuring effective feature transformation. Experiments on LibriSpeech and FLEURS datasets show MKL-VC significantly improves content preservation and robustness with short reference audio, outperforming kNN-VC. MKL-VC achieves performance comparable to FACodec, especially in cross-lingual voice conversion domain.
title Training-Free Voice Conversion with Factorized Optimal Transport
topic Sound
Computer Vision and Pattern Recognition
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.09709