FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Zeyu, Zhang, Yaoyun, Xu, Xuenan, Yin, Yongkang, Li, Chenxing, Wu, Mengyue, Zou, Yuexian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912607586222080
author Xie, Zeyu
Zhang, Yaoyun
Xu, Xuenan
Yin, Yongkang
Li, Chenxing
Wu, Mengyue
Zou, Yuexian
author_facet Xie, Zeyu
Zhang, Yaoyun
Xu, Xuenan
Yin, Yongkang
Li, Chenxing
Wu, Mengyue
Zou, Yuexian
contents The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have explored this task, existing methods largely focus on binary classification and fall short in explaining how manipulations occur, tracing where the sources originated, or generalizing to unseen sources-thereby limiting the explainability and reliability of detection. To address these limitations, we present FakeSound2, a benchmark designed to advance deepfake sound detection beyond binary accuracy. FakeSound2 evaluates models across three dimensions: localization, traceability, and generalization, covering 6 manipulation types and 12 diverse sources. Experimental results show that although current systems achieve high classification accuracy, they struggle to recognize forged pattern distributions and provide reliable explanations. By highlighting these gaps, FakeSound2 establishes a comprehensive benchmark that reveals key challenges and aims to foster robust, explainable, and generalizable approaches for trustworthy audio authentication.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17162
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
Xie, Zeyu
Zhang, Yaoyun
Xu, Xuenan
Yin, Yongkang
Li, Chenxing
Wu, Mengyue
Zou, Yuexian
Sound
Audio and Speech Processing
68Txx
I.2
The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have explored this task, existing methods largely focus on binary classification and fall short in explaining how manipulations occur, tracing where the sources originated, or generalizing to unseen sources-thereby limiting the explainability and reliability of detection. To address these limitations, we present FakeSound2, a benchmark designed to advance deepfake sound detection beyond binary accuracy. FakeSound2 evaluates models across three dimensions: localization, traceability, and generalization, covering 6 manipulation types and 12 diverse sources. Experimental results show that although current systems achieve high classification accuracy, they struggle to recognize forged pattern distributions and provide reliable explanations. By highlighting these gaps, FakeSound2 establishes a comprehensive benchmark that reveals key challenges and aims to foster robust, explainable, and generalizable approaches for trustworthy audio authentication.
title FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
topic Sound
Audio and Speech Processing
68Txx
I.2
url https://arxiv.org/abs/2509.17162