Quantifying Source Speaker Leakage in One-to-One Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wellington, Scott, Liu, Xuechen, Yamagishi, Junichi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916702147575808
author Wellington, Scott
Liu, Xuechen
Yamagishi, Junichi
author_facet Wellington, Scott
Liu, Xuechen
Yamagishi, Junichi
contents Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker's identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a "worst-case" white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers' data.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantifying Source Speaker Leakage in One-to-One Voice Conversion
Wellington, Scott
Liu, Xuechen
Yamagishi, Junichi
Sound
Cryptography and Security
Audio and Speech Processing
Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker's identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a "worst-case" white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers' data.
title Quantifying Source Speaker Leakage in One-to-One Voice Conversion
topic Sound
Cryptography and Security
Audio and Speech Processing
url https://arxiv.org/abs/2504.15822