A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grinberg, Petr, Kumar, Ankur, Koppisetti, Surya, Bharaj, Gaurav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914167818027008
author Grinberg, Petr
Kumar, Ankur
Koppisetti, Surya
Bharaj, Gaurav
author_facet Grinberg, Petr
Kumar, Ankur
Koppisetti, Surya
Bharaj, Gaurav
contents Evaluating explainability techniques, such as SHAP and LRP, in the context of audio deepfake detection is challenging due to lack of clear ground truth annotations. In the cases when we are able to obtain the ground truth, we find that these methods struggle to provide accurate explanations. In this work, we propose a novel data-driven approach to identify artifact regions in deepfake audio. We consider paired real and vocoded audio, and use the difference in time-frequency representation as the ground-truth explanation. The difference signal then serves as a supervision to train a diffusion model to expose the deepfake artifacts in a given vocoded audio. Experimental results on the VocV4 and LibriSeVoc datasets demonstrate that our method outperforms traditional explainability techniques, both qualitatively and quantitatively.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03425
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations
Grinberg, Petr
Kumar, Ankur
Koppisetti, Surya
Bharaj, Gaurav
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Evaluating explainability techniques, such as SHAP and LRP, in the context of audio deepfake detection is challenging due to lack of clear ground truth annotations. In the cases when we are able to obtain the ground truth, we find that these methods struggle to provide accurate explanations. In this work, we propose a novel data-driven approach to identify artifact regions in deepfake audio. We consider paired real and vocoded audio, and use the difference in time-frequency representation as the ground-truth explanation. The difference signal then serves as a supervision to train a diffusion model to expose the deepfake artifacts in a given vocoded audio. Experimental results on the VocV4 and LibriSeVoc datasets demonstrate that our method outperforms traditional explainability techniques, both qualitatively and quantitatively.
title A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.03425