SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Yi, Koppisetti, Surya, Tran, Trang, Bharaj, Gaurav
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914888293548032
author Zhu, Yi
Koppisetti, Surya
Tran, Trang
Bharaj, Gaurav
author_facet Zhu, Yi
Koppisetti, Surya
Tran, Trang
Bharaj, Gaurav
contents Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain data. Moreover, the black-box nature of existing models limits their use in real-world scenarios, where explanations are required for model decisions. To alleviate these issues, we introduce a new ADD model that explicitly uses the StyleLInguistics Mismatch (SLIM) in fake speech to separate them from real speech. SLIM first employs self-supervised pretraining on only real samples to learn the style-linguistics dependency in the real class. The learned features are then used in complement with standard pretrained acoustic features (e.g., Wav2vec) to learn a classifier on the real and fake classes. When the feature encoders are frozen, SLIM outperforms benchmark methods on out-of-domain datasets while achieving competitive results on in-domain data. The features learned by SLIM allow us to quantify the (mis)match between style and linguistic content in a sample, hence facilitating an explanation of the model decision.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18517
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
Zhu, Yi
Koppisetti, Surya
Tran, Trang
Bharaj, Gaurav
Sound
Artificial Intelligence
Audio and Speech Processing
Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain data. Moreover, the black-box nature of existing models limits their use in real-world scenarios, where explanations are required for model decisions. To alleviate these issues, we introduce a new ADD model that explicitly uses the StyleLInguistics Mismatch (SLIM) in fake speech to separate them from real speech. SLIM first employs self-supervised pretraining on only real samples to learn the style-linguistics dependency in the real class. The learned features are then used in complement with standard pretrained acoustic features (e.g., Wav2vec) to learn a classifier on the real and fake classes. When the feature encoders are frozen, SLIM outperforms benchmark methods on out-of-domain datasets while achieving competitive results on in-domain data. The features learned by SLIM allow us to quantify the (mis)match between style and linguistic content in a sample, hence facilitating an explanation of the model decision.
title SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2407.18517