Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kwok, Chin Yuen, Yip, Jia Qi, Qiu, Zhen, Chi, Chi Hung, Lam, Kwok Yan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909781491449856
author Kwok, Chin Yuen
Yip, Jia Qi
Qiu, Zhen
Chi, Chi Hung
Lam, Kwok Yan
author_facet Kwok, Chin Yuen
Yip, Jia Qi
Qiu, Zhen
Chi, Chi Hung
Lam, Kwok Yan
contents Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers with more samples, underrepresenting others and reducing the overall reliability of EER. Additionally, most ADD datasets lack diversity in bona fide speech, often featuring a single environment and speech style (e.g., clean read speech), limiting their ability to simulate real-world conditions. To address these challenges, we propose bona fide cross-testing, a novel evaluation framework that incorporates diverse bona fide datasets and aggregates EERs for more balanced assessments. Our approach improves robustness and interpretability compared to traditional evaluation methods. We benchmark over 150 synthesizers across nine bona fide speech types and release a new dataset to facilitate further research at https://github.com/cyaaronk/audio_deepfake_eval.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
Kwok, Chin Yuen
Yip, Jia Qi
Qiu, Zhen
Chi, Chi Hung
Lam, Kwok Yan
Sound
Artificial Intelligence
Computation and Language
Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers with more samples, underrepresenting others and reducing the overall reliability of EER. Additionally, most ADD datasets lack diversity in bona fide speech, often featuring a single environment and speech style (e.g., clean read speech), limiting their ability to simulate real-world conditions. To address these challenges, we propose bona fide cross-testing, a novel evaluation framework that incorporates diverse bona fide datasets and aggregates EERs for more balanced assessments. Our approach improves robustness and interpretability compared to traditional evaluation methods. We benchmark over 150 synthesizers across nine bona fide speech types and release a new dataset to facilitate further research at https://github.com/cyaaronk/audio_deepfake_eval.
title Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
topic Sound
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.09204