From Sharpness to Better Generalization for Speech Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wen, Liu, Xuechen, Wang, Xin, Yamagishi, Junichi, Qian, Yanmin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911004605022208
author Huang, Wen
Liu, Xuechen
Wang, Xin
Yamagishi, Junichi
Qian, Yanmin
author_facet Huang, Wen
Liu, Xuechen
Wang, Xin
Yamagishi, Junichi
Qian, Yanmin
contents Generalization remains a critical challenge in speech deepfake detection (SDD). While various approaches aim to improve robustness, generalization is typically assessed through performance metrics like equal error rate without a theoretical framework to explain model performance. This work investigates sharpness as a theoretical proxy for generalization in SDD. We analyze how sharpness responds to domain shifts and find it increases in unseen conditions, indicating higher model sensitivity. Based on this, we apply Sharpness-Aware Minimization (SAM) to reduce sharpness explicitly, leading to better and more stable performance across diverse unseen test sets. Furthermore, correlation analysis confirms a statistically significant relationship between sharpness and generalization in most test settings. These findings suggest that sharpness can serve as a theoretical indicator for generalization in SDD and that sharpness-aware training offers a promising strategy for improving robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Sharpness to Better Generalization for Speech Deepfake Detection
Huang, Wen
Liu, Xuechen
Wang, Xin
Yamagishi, Junichi
Qian, Yanmin
Audio and Speech Processing
Sound
Generalization remains a critical challenge in speech deepfake detection (SDD). While various approaches aim to improve robustness, generalization is typically assessed through performance metrics like equal error rate without a theoretical framework to explain model performance. This work investigates sharpness as a theoretical proxy for generalization in SDD. We analyze how sharpness responds to domain shifts and find it increases in unseen conditions, indicating higher model sensitivity. Based on this, we apply Sharpness-Aware Minimization (SAM) to reduce sharpness explicitly, leading to better and more stable performance across diverse unseen test sets. Furthermore, correlation analysis confirms a statistically significant relationship between sharpness and generalization in most test settings. These findings suggest that sharpness can serve as a theoretical indicator for generalization in SDD and that sharpness-aware training offers a promising strategy for improving robustness.
title From Sharpness to Better Generalization for Speech Deepfake Detection
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.11532