Reference-free Adversarial Sex Obfuscation in Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qu, Yangyang, Panariello, Michele, Todisco, Massimiliano, Evans, Nicholas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911090517999616
author Qu, Yangyang
Panariello, Michele
Todisco, Massimiliano
Evans, Nicholas
author_facet Qu, Yangyang
Panariello, Michele
Todisco, Massimiliano
Evans, Nicholas
contents Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02295
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reference-free Adversarial Sex Obfuscation in Speech
Qu, Yangyang
Panariello, Michele
Todisco, Massimiliano
Evans, Nicholas
Audio and Speech Processing
Sound
Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation.
title Reference-free Adversarial Sex Obfuscation in Speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2508.02295