Reference-free Adversarial Sex Obfuscation in Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911090517999616 |
|---|---|
| author | Qu, Yangyang Panariello, Michele Todisco, Massimiliano Evans, Nicholas |
| author_facet | Qu, Yangyang Panariello, Michele Todisco, Massimiliano Evans, Nicholas |
| contents | Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_02295 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Reference-free Adversarial Sex Obfuscation in Speech Qu, Yangyang Panariello, Michele Todisco, Massimiliano Evans, Nicholas Audio and Speech Processing Sound Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation. |
| title | Reference-free Adversarial Sex Obfuscation in Speech |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2508.02295 |