Continuous Learning of Transformer-based Audio Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Tuan Duy Nguyen, Teh, Kah Kuan, Tran, Huy Dat
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913494360653824
author Le, Tuan Duy Nguyen
Teh, Kah Kuan
Tran, Huy Dat
author_facet Le, Tuan Duy Nguyen
Teh, Kah Kuan
Tran, Huy Dat
contents This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a few-shot learning manner. Specifically, we conduct a large audio deepfake collection using various deep audio generation methods. The data is further enhanced with additional augmentation methods to increase variations amidst compressions, far-field recordings, noise, and other distortions. We then adopt the Audio Spectrogram Transformer for the audio deepfake detection model. Accordingly, the proposed method achieves promising performance on various benchmark datasets. Furthermore, we present a continuous learning plugin module to update the trained model most effectively with the fewest possible labeled data points of the new fake type. The proposed method outperforms the conventional direct fine-tuning approach with much fewer labeled data points.
format Preprint
id arxiv_https___arxiv_org_abs_2409_05924
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous Learning of Transformer-based Audio Deepfake Detection
Le, Tuan Duy Nguyen
Teh, Kah Kuan
Tran, Huy Dat
Sound
Audio and Speech Processing
This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a few-shot learning manner. Specifically, we conduct a large audio deepfake collection using various deep audio generation methods. The data is further enhanced with additional augmentation methods to increase variations amidst compressions, far-field recordings, noise, and other distortions. We then adopt the Audio Spectrogram Transformer for the audio deepfake detection model. Accordingly, the proposed method achieves promising performance on various benchmark datasets. Furthermore, we present a continuous learning plugin module to update the trained model most effectively with the fewest possible labeled data points of the new fake type. The proposed method outperforms the conventional direct fine-tuning approach with much fewer labeled data points.
title Continuous Learning of Transformer-based Audio Deepfake Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.05924