Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yi, Goel, Chirag, Koppisetti, Surya, Tran, Trang, Kumar, Ankur, Bharaj, Gaurav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914969173360640
author Zhu, Yi
Goel, Chirag
Koppisetti, Surya
Tran, Trang
Kumar, Ankur
Bharaj, Gaurav
author_facet Zhu, Yi
Goel, Chirag
Koppisetti, Surya
Tran, Trang
Kumar, Ankur
Bharaj, Gaurav
contents Audio deepfake detection is crucial to combat the malicious use of AI-synthesized speech. Among many efforts undertaken by the community, the ASVspoof challenge has become one of the benchmarks to evaluate the generalizability and robustness of detection models. In this paper, we present Reality Defender's submission to the ASVspoof5 challenge, highlighting a novel pretraining strategy which significantly improves generalizability while maintaining low computational cost during training. Our system SLIM learns the style-linguistics dependency embeddings from various types of bonafide speech using self-supervised contrastive learning. The learned embeddings help to discriminate spoof from bonafide speech by focusing on the relationship between the style and linguistics aspects. We evaluated our system on ASVspoof5, ASV2019, and In-the-wild. Our submission achieved minDCF of 0.1499 and EER of 5.5% on ASVspoof5 Track 1, and EER of 7.4% and 10.8% on ASV2019 and In-the-wild respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07379
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge
Zhu, Yi
Goel, Chirag
Koppisetti, Surya
Tran, Trang
Kumar, Ankur
Bharaj, Gaurav
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Audio deepfake detection is crucial to combat the malicious use of AI-synthesized speech. Among many efforts undertaken by the community, the ASVspoof challenge has become one of the benchmarks to evaluate the generalizability and robustness of detection models. In this paper, we present Reality Defender's submission to the ASVspoof5 challenge, highlighting a novel pretraining strategy which significantly improves generalizability while maintaining low computational cost during training. Our system SLIM learns the style-linguistics dependency embeddings from various types of bonafide speech using self-supervised contrastive learning. The learned embeddings help to discriminate spoof from bonafide speech by focusing on the relationship between the style and linguistics aspects. We evaluated our system on ASVspoof5, ASV2019, and In-the-wild. Our submission achieved minDCF of 0.1499 and EER of 5.5% on ASVspoof5 Track 1, and EER of 7.4% and 10.8% on ASV2019 and In-the-wild respectively.
title Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.07379