AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911548901949440 |
|---|---|
| author | Nguyen-Le, Hai-Son Nguyen-Thanh, Hung-Cuong Le-Khac, Nhien-An Nguyen, Dinh-Thuc Nguyen-Le, Hong-Hanh |
| author_facet | Nguyen-Le, Hai-Son Nguyen-Thanh, Hung-Cuong Le-Khac, Nhien-An Nguyen, Dinh-Thuc Nguyen-Le, Hong-Hanh |
| contents | The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_26856 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection Nguyen-Le, Hai-Son Nguyen-Thanh, Hung-Cuong Le-Khac, Nhien-An Nguyen, Dinh-Thuc Nguyen-Le, Hong-Hanh Sound Artificial Intelligence Audio and Speech Processing The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS. |
| title | AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2603.26856 |