Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Glazer, Neta, Chernin, David, Achituve, Idan, Gannot, Sharon, Fetaya, Ethan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918038282960896
author Glazer, Neta
Chernin, David
Achituve, Idan
Gannot, Sharon
Fetaya, Ethan
author_facet Glazer, Neta
Chernin, David
Achituve, Idan
Gannot, Sharon
Fetaya, Ethan
contents Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS systems continue to evolve, detection models must be able to efficiently adapt to previously unseen generation models with minimal data. This paper introduces ADD-GP, a few-shot adaptive framework based on a Gaussian Process (GP) classifier for Audio Deepfake Detection (ADD). We show how the combination of a powerful deep embedding model with the Gaussian processes flexibility can achieve strong performance and adaptability. Additionally, we show this approach can also be used for personalized detection, with greater robustness to new TTS models and one-shot adaptability. To support our evaluation, a benchmark dataset is constructed for this task using new state-of-the-art voice cloning models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
Glazer, Neta
Chernin, David
Achituve, Idan
Gannot, Sharon
Fetaya, Ethan
Sound
Machine Learning
Audio and Speech Processing
Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS systems continue to evolve, detection models must be able to efficiently adapt to previously unseen generation models with minimal data. This paper introduces ADD-GP, a few-shot adaptive framework based on a Gaussian Process (GP) classifier for Audio Deepfake Detection (ADD). We show how the combination of a powerful deep embedding model with the Gaussian processes flexibility can achieve strong performance and adaptability. Additionally, we show this approach can also be used for personalized detection, with greater robustness to new TTS models and one-shot adaptability. To support our evaluation, a benchmark dataset is constructed for this task using new state-of-the-art voice cloning models.
title Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2505.23619