The First Environmental Sound Deepfake Detection Challenge: Benchmarking Robustness, Evaluation, and Insights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Han, Xiao, Yang, Das, Rohan Kumar, Bai, Jisheng, Dang, Ting
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912950530342912
author Yin, Han
Xiao, Yang
Das, Rohan Kumar
Bai, Jisheng
Dang, Ting
author_facet Yin, Han
Xiao, Yang
Das, Rohan Kumar
Bai, Jisheng
Dang, Ting
contents Recent progress in audio generation has made it increasingly easy to create highly realistic environmental soundscapes, which can be misused to produce deceptive content, such as fake alarms, gunshots, and crowd sounds, raising concerns for public safety and trust. While deepfake detection for speech and singing voice has been extensively studied, environmental sound deepfake detection (ESDD) remains underexplored. To advance ESDD, the first edition of the ESDD challenge was launched, attracting 97 registered teams and receiving 1,748 valid submissions. This paper presents the task formulation, dataset construction, evaluation protocols, baseline systems, and key insights from the challenge results. Furthermore, we analyze common architectural choices and training strategies among top-performing systems. Finally, we discuss potential future research directions for ESDD, outlining key opportunities and open problems to guide subsequent studies in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04865
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The First Environmental Sound Deepfake Detection Challenge: Benchmarking Robustness, Evaluation, and Insights
Yin, Han
Xiao, Yang
Das, Rohan Kumar
Bai, Jisheng
Dang, Ting
Sound
Recent progress in audio generation has made it increasingly easy to create highly realistic environmental soundscapes, which can be misused to produce deceptive content, such as fake alarms, gunshots, and crowd sounds, raising concerns for public safety and trust. While deepfake detection for speech and singing voice has been extensively studied, environmental sound deepfake detection (ESDD) remains underexplored. To advance ESDD, the first edition of the ESDD challenge was launched, attracting 97 registered teams and receiving 1,748 valid submissions. This paper presents the task formulation, dataset construction, evaluation protocols, baseline systems, and key insights from the challenge results. Furthermore, we analyze common architectural choices and training strategies among top-performing systems. Finally, we discuss potential future research directions for ESDD, outlining key opportunities and open problems to guide subsequent studies in this field.
title The First Environmental Sound Deepfake Detection Challenge: Benchmarking Robustness, Evaluation, and Insights
topic Sound
url https://arxiv.org/abs/2603.04865