Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Coletta, Emma, Salvi, Davide, Negroni, Viola, Leonzio, Daniele Ugo, Bestagini, Paolo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912289595064320
author Coletta, Emma
Salvi, Davide
Negroni, Viola
Leonzio, Daniele Ugo
Bestagini, Paolo
author_facet Coletta, Emma
Salvi, Davide
Negroni, Viola
Leonzio, Daniele Ugo
Bestagini, Paolo
contents The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting speech deepfakes primarily rely on supervised learning, which suffers from two critical limitations: limited generalization to unseen synthesis techniques and a lack of explainability. In this paper, we address these issues by introducing a novel interpretable one-class detection framework, which reframes speech deepfake detection as an anomaly detection task. Our model is trained exclusively on real speech to characterize its distribution, enabling the classification of out-of-distribution samples as synthetically generated. Additionally, our framework produces interpretable anomaly maps during inference, highlighting anomalous regions across both time and frequency domains. This is done through a Student-Teacher Feature Pyramid Matching system, enhanced with Discrepancy Scaling to improve generalization capabilities across unseen data distributions. Extensive evaluations demonstrate the superior performance of our approach compared to the considered baselines, validating the effectiveness of framing speech deepfake detection as an anomaly detection problem.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching
Coletta, Emma
Salvi, Davide
Negroni, Viola
Leonzio, Daniele Ugo
Bestagini, Paolo
Sound
Computer Vision and Pattern Recognition
Multimedia
The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting speech deepfakes primarily rely on supervised learning, which suffers from two critical limitations: limited generalization to unseen synthesis techniques and a lack of explainability. In this paper, we address these issues by introducing a novel interpretable one-class detection framework, which reframes speech deepfake detection as an anomaly detection task. Our model is trained exclusively on real speech to characterize its distribution, enabling the classification of out-of-distribution samples as synthetically generated. Additionally, our framework produces interpretable anomaly maps during inference, highlighting anomalous regions across both time and frequency domains. This is done through a Student-Teacher Feature Pyramid Matching system, enhanced with Discrepancy Scaling to improve generalization capabilities across unseen data distributions. Extensive evaluations demonstrate the superior performance of our approach compared to the considered baselines, validating the effectiveness of framing speech deepfake detection as an anomaly detection problem.
title Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching
topic Sound
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2503.18032