SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, You, Zang, Yongyi, Shi, Jiatong, Yamamoto, Ryuichi, Toda, Tomoki, Duan, Zhiyao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918303928156160
author Zhang, You
Zang, Yongyi
Shi, Jiatong
Yamamoto, Ryuichi
Toda, Tomoki
Duan, Zhiyao
author_facet Zhang, You
Zang, Yongyi
Shi, Jiatong
Yamamoto, Ryuichi
Toda, Tomoki
Duan, Zhiyao
contents With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-generated singing voices from authentic singers. This challenge features two tracks: a controlled setting track (CtrSVDD) and an in-the-wild scenario track (WildSVDD). The CtrSVDD track utilizes publicly available singing vocal data to generate deepfakes using state-of-the-art singing voice synthesis and conversion systems. Meanwhile, the WildSVDD track expands upon the existing SingFake dataset, which includes data sourced from popular user-generated content websites. For the CtrSVDD track, we received submissions from 47 teams, with 37 surpassing our baselines and the top team achieving a 1.65% equal error rate. For the WildSVDD track, we benchmarked the baselines. This paper reviews these results, discusses key findings, and outlines future directions for SVDD research.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16132
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
Zhang, You
Zang, Yongyi
Shi, Jiatong
Yamamoto, Ryuichi
Toda, Tomoki
Duan, Zhiyao
Audio and Speech Processing
Multimedia
Sound
With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-generated singing voices from authentic singers. This challenge features two tracks: a controlled setting track (CtrSVDD) and an in-the-wild scenario track (WildSVDD). The CtrSVDD track utilizes publicly available singing vocal data to generate deepfakes using state-of-the-art singing voice synthesis and conversion systems. Meanwhile, the WildSVDD track expands upon the existing SingFake dataset, which includes data sourced from popular user-generated content websites. For the CtrSVDD track, we received submissions from 47 teams, with 37 surpassing our baselines and the top team achieving a 1.65% equal error rate. For the WildSVDD track, we benchmarked the baselines. This paper reviews these results, discusses key findings, and outlines future directions for SVDD research.
title SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
topic Audio and Speech Processing
Multimedia
Sound
url https://arxiv.org/abs/2408.16132