Towards generalisable and calibrated synthetic speech detection with self-supervised representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pascu, Octavian, Stan, Adriana, Oneata, Dan, Oneata, Elisabeta, Cucu, Horia
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913389009174528
author Pascu, Octavian
Stan, Adriana
Oneata, Dan
Oneata, Elisabeta
Cucu, Horia
author_facet Pascu, Octavian
Stan, Adriana
Oneata, Dan
Oneata, Elisabeta
Cucu, Horia
contents Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deepfake detectors. However, recent studies have shown that the current audio deepfake models fall short of this desideratum. In this work we investigate the potential of pretrained self-supervised representations in building general and calibrated audio deepfake detection models. We show that large frozen representations coupled with a simple logistic regression classifier are extremely effective in achieving strong generalisation capabilities: compared to the RawNet2 model, this approach reduces the equal error rate from 30.9% to 8.8% on a benchmark of eight deepfake datasets, while learning less than 2k parameters. Moreover, the proposed method produces considerably more reliable predictions compared to previous approaches making it more suitable for realistic use.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05384
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards generalisable and calibrated synthetic speech detection with self-supervised representations
Pascu, Octavian
Stan, Adriana
Oneata, Dan
Oneata, Elisabeta
Cucu, Horia
Audio and Speech Processing
Sound
Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deepfake detectors. However, recent studies have shown that the current audio deepfake models fall short of this desideratum. In this work we investigate the potential of pretrained self-supervised representations in building general and calibrated audio deepfake detection models. We show that large frozen representations coupled with a simple logistic regression classifier are extremely effective in achieving strong generalisation capabilities: compared to the RawNet2 model, this approach reduces the equal error rate from 30.9% to 8.8% on a benchmark of eight deepfake datasets, while learning less than 2k parameters. Moreover, the proposed method produces considerably more reliable predictions compared to previous approaches making it more suitable for realistic use.
title Towards generalisable and calibrated synthetic speech detection with self-supervised representations
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.05384