AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen-Le, Hai-Son, Nguyen-Thanh, Hung-Cuong, Le-Khac, Nhien-An, Nguyen, Dinh-Thuc, Nguyen-Le, Hong-Hanh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911548901949440
author Nguyen-Le, Hai-Son
Nguyen-Thanh, Hung-Cuong
Le-Khac, Nhien-An
Nguyen, Dinh-Thuc
Nguyen-Le, Hong-Hanh
author_facet Nguyen-Le, Hai-Son
Nguyen-Thanh, Hung-Cuong
Le-Khac, Nhien-An
Nguyen, Dinh-Thuc
Nguyen-Le, Hong-Hanh
contents The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26856
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
Nguyen-Le, Hai-Son
Nguyen-Thanh, Hung-Cuong
Le-Khac, Nhien-An
Nguyen, Dinh-Thuc
Nguyen-Le, Hong-Hanh
Sound
Artificial Intelligence
Audio and Speech Processing
The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS.
title AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2603.26856