Mitigating Unauthorized Speech Synthesis for Voice Protection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhisheng, Yang, Qianyi, Wang, Derui, Huang, Pengyang, Cao, Yuxin, Ye, Kai, Hao, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912089517326336
author Zhang, Zhisheng
Yang, Qianyi
Wang, Derui
Huang, Pengyang
Cao, Yuxin
Ye, Kai
Hao, Jie
author_facet Zhang, Zhisheng
Yang, Qianyi
Wang, Derui
Huang, Pengyang
Cao, Yuxin
Ye, Kai
Hao, Jie
contents With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our daily lives. Therefore, it is crucial to protect publicly accessible speech data that contains sensitive information, such as personal voiceprints. Most previous defense methods have focused on spoofing speaker verification systems in timbre similarity but the synthesized deepfake speech is still of high quality. In response to the rising hazards, we devise an effective, transferable, and robust proactive protection technology named Pivotal Objective Perturbation (POP) that applies imperceptible error-minimizing noises on original speech samples to prevent them from being effectively learned for text-to-speech (TTS) synthesis models so that high-quality deepfake speeches cannot be generated. We conduct extensive experiments on state-of-the-art (SOTA) TTS models utilizing objective and subjective metrics to comprehensively evaluate our proposed method. The experimental results demonstrate outstanding effectiveness and transferability across various models. Compared to the speech unclarity score of 21.94% from voice synthesizers trained on samples without protection, POP-protected samples significantly increase it to 127.31%. Moreover, our method shows robustness against noise reduction and data augmentation techniques, thereby greatly reducing potential hazards.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20742
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mitigating Unauthorized Speech Synthesis for Voice Protection
Zhang, Zhisheng
Yang, Qianyi
Wang, Derui
Huang, Pengyang
Cao, Yuxin
Ye, Kai
Hao, Jie
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our daily lives. Therefore, it is crucial to protect publicly accessible speech data that contains sensitive information, such as personal voiceprints. Most previous defense methods have focused on spoofing speaker verification systems in timbre similarity but the synthesized deepfake speech is still of high quality. In response to the rising hazards, we devise an effective, transferable, and robust proactive protection technology named Pivotal Objective Perturbation (POP) that applies imperceptible error-minimizing noises on original speech samples to prevent them from being effectively learned for text-to-speech (TTS) synthesis models so that high-quality deepfake speeches cannot be generated. We conduct extensive experiments on state-of-the-art (SOTA) TTS models utilizing objective and subjective metrics to comprehensively evaluate our proposed method. The experimental results demonstrate outstanding effectiveness and transferability across various models. Compared to the speech unclarity score of 21.94% from voice synthesizers trained on samples without protection, POP-protected samples significantly increase it to 127.31%. Moreover, our method shows robustness against noise reduction and data augmentation techniques, thereby greatly reducing potential hazards.
title Mitigating Unauthorized Speech Synthesis for Voice Protection
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.20742