Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kukanov, Ivan, Laakkonen, Janne, Kinnunen, Tomi, Hautamäki, Ville
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915000412536832
author Kukanov, Ivan
Laakkonen, Janne
Kinnunen, Tomi
Hautamäki, Ville
author_facet Kukanov, Ivan
Laakkonen, Janne
Kinnunen, Tomi
Hautamäki, Ville
contents Current speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need for systems that can generalize to unseen attacks not observed during training. We address this problem from the perspective of meta-learning, aiming to learn attack-invariant features to adapt to unseen attacks with very few samples available. This approach is promising since generating of a high-scale training dataset is often expensive or infeasible. Our experiments demonstrated an improvement in the Equal Error Rate (EER) from 21.67% to 10.42% on the InTheWild dataset, using just 96 samples from the unseen dataset. Continuous few-shot adaptation ensures that the system remains up-to-date.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20578
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes
Kukanov, Ivan
Laakkonen, Janne
Kinnunen, Tomi
Hautamäki, Ville
Audio and Speech Processing
Artificial Intelligence
Sound
Current speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need for systems that can generalize to unseen attacks not observed during training. We address this problem from the perspective of meta-learning, aiming to learn attack-invariant features to adapt to unseen attacks with very few samples available. This approach is promising since generating of a high-scale training dataset is often expensive or infeasible. Our experiments demonstrated an improvement in the Equal Error Rate (EER) from 21.67% to 10.42% on the InTheWild dataset, using just 96 samples from the unseen dataset. Continuous few-shot adaptation ensures that the system remains up-to-date.
title Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2410.20578