ESDD 2026: Environmental Sound Deepfake Detection Challenge Evaluation Plan

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Han, Xiao, Yang, Das, Rohan Kumar, Bai, Jisheng, Dang, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912787160104960
author Yin, Han
Xiao, Yang
Das, Rohan Kumar
Bai, Jisheng
Dang, Ting
author_facet Yin, Han
Xiao, Yang
Das, Rohan Kumar
Bai, Jisheng
Dang, Ting
contents Recent advances in audio generation systems have enabled the creation of highly realistic and immersive soundscapes, which are increasingly used in film and virtual reality. However, these audio generators also raise concerns about potential misuse, such as generating deceptive audio content for fake videos and spreading misleading information. Existing datasets for environmental sound deepfake detection (ESDD) are limited in scale and audio types. To address this gap, we have proposed EnvSDD, the first large-scale curated dataset designed for ESDD, consisting of 45.25 hours of real and 316.7 hours of fake sound. Based on EnvSDD, we are launching the Environmental Sound Deepfake Detection Challenge. Specifically, we present two different tracks: ESDD in Unseen Generators and Black-Box Low-Resource ESDD, covering various challenges encountered in real-life scenarios. The challenge will be held in conjunction with the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).
format Preprint
id arxiv_https___arxiv_org_abs_2508_04529
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ESDD 2026: Environmental Sound Deepfake Detection Challenge Evaluation Plan
Yin, Han
Xiao, Yang
Das, Rohan Kumar
Bai, Jisheng
Dang, Ting
Sound
Recent advances in audio generation systems have enabled the creation of highly realistic and immersive soundscapes, which are increasingly used in film and virtual reality. However, these audio generators also raise concerns about potential misuse, such as generating deceptive audio content for fake videos and spreading misleading information. Existing datasets for environmental sound deepfake detection (ESDD) are limited in scale and audio types. To address this gap, we have proposed EnvSDD, the first large-scale curated dataset designed for ESDD, consisting of 45.25 hours of real and 316.7 hours of fake sound. Based on EnvSDD, we are launching the Environmental Sound Deepfake Detection Challenge. Specifically, we present two different tracks: ESDD in Unseen Generators and Black-Box Low-Resource ESDD, covering various challenges encountered in real-life scenarios. The challenge will be held in conjunction with the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).
title ESDD 2026: Environmental Sound Deepfake Detection Challenge Evaluation Plan
topic Sound
url https://arxiv.org/abs/2508.04529