Discrete Optimal Transport and Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Selitskiy, Anton, Kocharekar, Maitreya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917297680023552
author Selitskiy, Anton
Kocharekar, Maitreya
author_facet Selitskiy, Anton
Kocharekar, Maitreya
contents In this work, we address the task of voice conversion (VC) using a vector-based interface. To align audio embeddings across speakers, we employ discrete optimal transport (OT) and approximate the transport map using the barycentric projection. Our evaluation demonstrates that this approach yields high-quality and effective voice conversion. We also perform an ablation study on the number of embeddings used, extending previous work on simple averaging of kNN and OT results. Additionally, we show that applying discrete OT as a post-processing step in audio generation can cause synthetic speech to be misclassified as real, revealing a novel and strong adversarial attack.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04382
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discrete Optimal Transport and Voice Conversion
Selitskiy, Anton
Kocharekar, Maitreya
Audio and Speech Processing
Machine Learning
Sound
In this work, we address the task of voice conversion (VC) using a vector-based interface. To align audio embeddings across speakers, we employ discrete optimal transport (OT) and approximate the transport map using the barycentric projection. Our evaluation demonstrates that this approach yields high-quality and effective voice conversion. We also perform an ablation study on the number of embeddings used, extending previous work on simple averaging of kNN and OT results. Additionally, we show that applying discrete OT as a post-processing step in audio generation can cause synthetic speech to be misclassified as real, revealing a novel and strong adversarial attack.
title Discrete Optimal Transport and Voice Conversion
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2505.04382