What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Grinberg, Petr, Kumar, Ankur, Koppisetti, Surya, Bharaj, Gaurav
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910996472266752
author Grinberg, Petr
Kumar, Ankur
Koppisetti, Surya
Bharaj, Gaurav
author_facet Grinberg, Petr
Kumar, Ankur
Koppisetti, Surya
Bharaj, Gaurav
contents Adding explanations to audio deepfake detection (ADD) models will boost their real-world application by providing insight on the decision making process. In this paper, we propose a relevancy-based explainable AI (XAI) method to analyze the predictions of transformer-based ADD models. We compare against standard Grad-CAM and SHAP-based methods, using quantitative faithfulness metrics as well as a partial spoof test, to comprehensively analyze the relative importance of different temporal regions in an audio. We consider large datasets, unlike previous works where only limited utterances are studied, and find that the XAI methods differ in their explanations. The proposed relevancy-based XAI method performs the best overall on a variety of metrics. Further investigation on the relative importance of speech/non-speech, phonetic content, and voice onsets/offsets suggest that the XAI results obtained from analyzing limited utterances don't necessarily hold when evaluated on large datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain
Grinberg, Petr
Kumar, Ankur
Koppisetti, Surya
Bharaj, Gaurav
Machine Learning
Sound
Audio and Speech Processing
Adding explanations to audio deepfake detection (ADD) models will boost their real-world application by providing insight on the decision making process. In this paper, we propose a relevancy-based explainable AI (XAI) method to analyze the predictions of transformer-based ADD models. We compare against standard Grad-CAM and SHAP-based methods, using quantitative faithfulness metrics as well as a partial spoof test, to comprehensively analyze the relative importance of different temporal regions in an audio. We consider large datasets, unlike previous works where only limited utterances are studied, and find that the XAI methods differ in their explanations. The proposed relevancy-based XAI method performs the best overall on a variety of metrics. Further investigation on the relative importance of speech/non-speech, phonetic content, and voice onsets/offsets suggest that the XAI results obtained from analyzing limited utterances don't necessarily hold when evaluated on large datasets.
title What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.13887