RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fu, Ruibo, Wang, Xiaopeng, Wen, Zhengqi, Tao, Jianhua, Xie, Yuankun, Wang, Zhiyong, Qiang, Chunyu, Liu, Xuefei, Fan, Cunhang, Li, Chenxing, Li, Guanjun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916770465447936
author Fu, Ruibo
Wang, Xiaopeng
Wen, Zhengqi
Tao, Jianhua
Xie, Yuankun
Wang, Zhiyong
Qiang, Chunyu
Liu, Xuefei
Fan, Cunhang
Li, Chenxing
Li, Guanjun
author_facet Fu, Ruibo
Wang, Xiaopeng
Wen, Zhengqi
Tao, Jianhua
Xie, Yuankun
Wang, Zhiyong
Qiang, Chunyu
Liu, Xuefei
Fan, Cunhang
Li, Chenxing
Li, Guanjun
contents Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attack patterns. This limitation mainly arises because the models rely heavily on the distribution of the training data and fail to learn a decision boundary that captures the essential characteristics of forgeries. Additionally, relying solely on a classification loss makes it difficult to capture the intrinsic differences between real and fake audio. In this paper, we propose the RPRA-ADD, an integrated Reconstruction-Perception-Reinforcement-Attention networks based forgery trace enhancement-driven robust audio deepfake detection framework. First, we propose a Global-Local Forgery Perception (GLFP) module for enhancing the acoustic perception capacity of forgery traces. To significantly reinforce the feature space distribution differences between real and fake audio, the Multi-stage Dispersed Enhancement Loss (MDEL) is designed, which implements a dispersal strategy in multi-stage feature spaces. Furthermore, in order to enhance feature awareness towards forgery traces, the Fake Trace Focused Attention (FTFA) mechanism is introduced to adjust attention weights dynamically according to the reconstruction discrepancy matrix. Visualization experiments not only demonstrate that FTFA improves attention to voice segments, but also enhance the generalization capability. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on 4 benchmark datasets, including ASVspoof2019, ASVspoof2021, CodecFake, and FakeSound, achieving over 20% performance improvement. In addition, it outperforms existing methods in rigorous 3*3 cross-domain evaluations across Speech, Sound, and Singing, demonstrating strong generalization capability across diverse audio domains.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00375
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
Fu, Ruibo
Wang, Xiaopeng
Wen, Zhengqi
Tao, Jianhua
Xie, Yuankun
Wang, Zhiyong
Qiang, Chunyu
Liu, Xuefei
Fan, Cunhang
Li, Chenxing
Li, Guanjun
Sound
Audio and Speech Processing
Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attack patterns. This limitation mainly arises because the models rely heavily on the distribution of the training data and fail to learn a decision boundary that captures the essential characteristics of forgeries. Additionally, relying solely on a classification loss makes it difficult to capture the intrinsic differences between real and fake audio. In this paper, we propose the RPRA-ADD, an integrated Reconstruction-Perception-Reinforcement-Attention networks based forgery trace enhancement-driven robust audio deepfake detection framework. First, we propose a Global-Local Forgery Perception (GLFP) module for enhancing the acoustic perception capacity of forgery traces. To significantly reinforce the feature space distribution differences between real and fake audio, the Multi-stage Dispersed Enhancement Loss (MDEL) is designed, which implements a dispersal strategy in multi-stage feature spaces. Furthermore, in order to enhance feature awareness towards forgery traces, the Fake Trace Focused Attention (FTFA) mechanism is introduced to adjust attention weights dynamically according to the reconstruction discrepancy matrix. Visualization experiments not only demonstrate that FTFA improves attention to voice segments, but also enhance the generalization capability. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on 4 benchmark datasets, including ASVspoof2019, ASVspoof2021, CodecFake, and FakeSound, achieving over 20% performance improvement. In addition, it outperforms existing methods in rigorous 3*3 cross-domain evaluations across Speech, Sound, and Singing, demonstrating strong generalization capability across diverse audio domains.
title RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00375