Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pianese, Alessandro, Cozzolino, Davide, Poggi, Giovanni, Verdoliva, Luisa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914853905498112
author Pianese, Alessandro
Cozzolino, Davide
Poggi, Giovanni
Verdoliva, Luisa
author_facet Pianese, Alessandro
Cozzolino, Davide
Poggi, Giovanni
Verdoliva, Luisa
contents Generalization is a main issue for current audio deepfake detectors, which struggle to provide reliable results on out-of-distribution data. Given the speed at which more and more accurate synthesis methods are developed, it is very important to design techniques that work well also on data they were not trained for. In this paper we study the potential of large-scale pre-trained models for audio deepfake detection, with special focus on generalization ability. To this end, the detection problem is reformulated in a speaker verification framework and fake audios are exposed by the mismatch between the voice sample under test and the voice of the claimed identity. With this paradigm, no fake speech sample is necessary in training, cutting off any link with the generation method at the root, and ensuring full generalization ability. Features are extracted by general-purpose large pre-trained models, with no need for training or fine-tuning on specific fake detection or speaker verification datasets. At detection time only a limited set of voice fragments of the identity under test is required. Experiments on several datasets widespread in the community show that detectors based on pre-trained models achieve excellent performance and show strong generalization ability, rivaling supervised methods on in-distribution data and largely overcoming them on out-of-distribution data.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
Pianese, Alessandro
Cozzolino, Davide
Poggi, Giovanni
Verdoliva, Luisa
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Generalization is a main issue for current audio deepfake detectors, which struggle to provide reliable results on out-of-distribution data. Given the speed at which more and more accurate synthesis methods are developed, it is very important to design techniques that work well also on data they were not trained for. In this paper we study the potential of large-scale pre-trained models for audio deepfake detection, with special focus on generalization ability. To this end, the detection problem is reformulated in a speaker verification framework and fake audios are exposed by the mismatch between the voice sample under test and the voice of the claimed identity. With this paradigm, no fake speech sample is necessary in training, cutting off any link with the generation method at the root, and ensuring full generalization ability. Features are extracted by general-purpose large pre-trained models, with no need for training or fine-tuning on specific fake detection or speaker verification datasets. At detection time only a limited set of voice fragments of the identity under test is required. Experiments on several datasets widespread in the community show that detectors based on pre-trained models achieve excellent performance and show strong generalization ability, rivaling supervised methods on in-distribution data and largely overcoming them on out-of-distribution data.
title Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2405.02179