I Can Hear You: Selective Robust Training for Deepfake Audio Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zirui, Hao, Wei, Sankoh, Aroon, Lin, William, Mendiola-Ortiz, Emanuel, Yang, Junfeng, Mao, Chengzhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912099617210368
author Zhang, Zirui
Hao, Wei
Sankoh, Aroon
Lin, William
Mendiola-Ortiz, Emanuel
Yang, Junfeng
Mao, Chengzhi
author_facet Zhang, Zirui
Hao, Wei
Sankoh, Aroon
Lin, William
Mendiola-Ortiz, Emanuel
Yang, Junfeng
Mao, Chengzhi
contents Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we establish the largest public voice dataset to date, named DeepFakeVox-HQ, comprising 1.3 million samples, including 270,000 high-quality deepfake samples from 14 diverse sources. Despite previously reported high accuracy, existing deepfake voice detectors struggle with our diversely collected dataset, and their detection success rates drop even further under realistic corruptions and adversarial attacks. We conduct a holistic investigation into factors that enhance model robustness and show that incorporating a diversified set of voice augmentations is beneficial. Moreover, we find that the best detection models often rely on high-frequency features, which are imperceptible to humans and can be easily manipulated by an attacker. To address this, we propose the F-SAT: Frequency-Selective Adversarial Training method focusing on high-frequency components. Empirical results demonstrate that using our training dataset boosts baseline model performance (without robust training) by 33%, and our robust training further improves accuracy by 7.7% on clean samples and by 29.3% on corrupted and attacked samples, over the state-of-the-art RawNet3 model.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00121
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle I Can Hear You: Selective Robust Training for Deepfake Audio Detection
Zhang, Zirui
Hao, Wei
Sankoh, Aroon
Lin, William
Mendiola-Ortiz, Emanuel
Yang, Junfeng
Mao, Chengzhi
Sound
Artificial Intelligence
Audio and Speech Processing
Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we establish the largest public voice dataset to date, named DeepFakeVox-HQ, comprising 1.3 million samples, including 270,000 high-quality deepfake samples from 14 diverse sources. Despite previously reported high accuracy, existing deepfake voice detectors struggle with our diversely collected dataset, and their detection success rates drop even further under realistic corruptions and adversarial attacks. We conduct a holistic investigation into factors that enhance model robustness and show that incorporating a diversified set of voice augmentations is beneficial. Moreover, we find that the best detection models often rely on high-frequency features, which are imperceptible to humans and can be easily manipulated by an attacker. To address this, we propose the F-SAT: Frequency-Selective Adversarial Training method focusing on high-frequency components. Empirical results demonstrate that using our training dataset boosts baseline model performance (without robust training) by 33%, and our robust training further improves accuracy by 7.7% on clean samples and by 29.3% on corrupted and attacked samples, over the state-of-the-art RawNet3 model.
title I Can Hear You: Selective Robust Training for Deepfake Audio Detection
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2411.00121