FairSSD: Understanding Bias in Synthetic Speech Detectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yadav, Amit Kumar Singh, Bhagtani, Kratika, Salvi, Davide, Bestagini, Paolo, Delp, Edward J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909172507869184
author Yadav, Amit Kumar Singh
Bhagtani, Kratika
Salvi, Davide
Bestagini, Paolo
Delp, Edward J.
author_facet Yadav, Amit Kumar Singh
Bhagtani, Kratika
Salvi, Davide
Bestagini, Paolo
Delp, Edward J.
contents Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of synthetic speech generated from these methods to commit fraud. To counter such misuse, many methods have been proposed to detect synthetic speech. Some of these detectors are more interpretable, can generalize to detect synthetic speech in the wild and are robust to noise. However, limited work has been done on understanding bias in these detectors. In this work, we examine bias in existing synthetic speech detectors to determine if they will unfairly target a particular gender, age and accent group. We also inspect whether these detectors will have a higher misclassification rate for bona fide speech from speech-impaired speakers w.r.t fluent speakers. Extensive experiments on 6 existing synthetic speech detectors using more than 0.9 million speech signals demonstrate that most detectors are gender, age and accent biased, and future work is needed to ensure fairness. To support future research, we release our evaluation dataset, models used in our study and source code at https://gitlab.com/viper-purdue/fairssd.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10989
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FairSSD: Understanding Bias in Synthetic Speech Detectors
Yadav, Amit Kumar Singh
Bhagtani, Kratika
Salvi, Davide
Bestagini, Paolo
Delp, Edward J.
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Sound
Audio and Speech Processing
Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of synthetic speech generated from these methods to commit fraud. To counter such misuse, many methods have been proposed to detect synthetic speech. Some of these detectors are more interpretable, can generalize to detect synthetic speech in the wild and are robust to noise. However, limited work has been done on understanding bias in these detectors. In this work, we examine bias in existing synthetic speech detectors to determine if they will unfairly target a particular gender, age and accent group. We also inspect whether these detectors will have a higher misclassification rate for bona fide speech from speech-impaired speakers w.r.t fluent speakers. Extensive experiments on 6 existing synthetic speech detectors using more than 0.9 million speech signals demonstrate that most detectors are gender, age and accent biased, and future work is needed to ensure fairness. To support future research, we release our evaluation dataset, models used in our study and source code at https://gitlab.com/viper-purdue/fairssd.
title FairSSD: Understanding Bias in Synthetic Speech Detectors
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2404.10989