Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915000412536832 |
|---|---|
| author | Kukanov, Ivan Laakkonen, Janne Kinnunen, Tomi Hautamäki, Ville |
| author_facet | Kukanov, Ivan Laakkonen, Janne Kinnunen, Tomi Hautamäki, Ville |
| contents | Current speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need for systems that can generalize to unseen attacks not observed during training. We address this problem from the perspective of meta-learning, aiming to learn attack-invariant features to adapt to unseen attacks with very few samples available. This approach is promising since generating of a high-scale training dataset is often expensive or infeasible. Our experiments demonstrated an improvement in the Equal Error Rate (EER) from 21.67% to 10.42% on the InTheWild dataset, using just 96 samples from the unseen dataset. Continuous few-shot adaptation ensures that the system remains up-to-date. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_20578 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes Kukanov, Ivan Laakkonen, Janne Kinnunen, Tomi Hautamäki, Ville Audio and Speech Processing Artificial Intelligence Sound Current speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need for systems that can generalize to unseen attacks not observed during training. We address this problem from the perspective of meta-learning, aiming to learn attack-invariant features to adapt to unseen attacks with very few samples available. This approach is promising since generating of a high-scale training dataset is often expensive or infeasible. Our experiments demonstrated an improvement in the Equal Error Rate (EER) from 21.67% to 10.42% on the InTheWild dataset, using just 96 samples from the unseen dataset. Continuous few-shot adaptation ensures that the system remains up-to-date. |
| title | Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes |
| topic | Audio and Speech Processing Artificial Intelligence Sound |
| url | https://arxiv.org/abs/2410.20578 |