Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yuankun, Fu, Ruibo, Wang, Zhiyong, Wang, Xiaopeng, Cao, Songjun, Ma, Long, Cheng, Haonan, Ye, Long
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915717632229376
author Xie, Yuankun
Fu, Ruibo
Wang, Zhiyong
Wang, Xiaopeng
Cao, Songjun
Ma, Long
Cheng, Haonan
Ye, Long
author_facet Xie, Yuankun
Fu, Ruibo
Wang, Zhiyong
Wang, Xiaopeng
Cao, Songjun
Ma, Long
Cheng, Haonan
Ye, Long
contents The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the all-type ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL front-end by learning specialized prompt tokens for ADD, requiring 458x fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types, we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception
Xie, Yuankun
Fu, Ruibo
Wang, Zhiyong
Wang, Xiaopeng
Cao, Songjun
Ma, Long
Cheng, Haonan
Ye, Long
Sound
Artificial Intelligence
The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the all-type ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL front-end by learning specialized prompt tokens for ADD, requiring 458x fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types, we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets.
title Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2504.06753