Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Kuiyuan, Hua, Zhongyun, Lan, Rushi, Zhang, Yushu, Guo, Yifang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915068824780800
author Zhang, Kuiyuan
Hua, Zhongyun
Lan, Rushi
Zhang, Yushu
Guo, Yifang
author_facet Zhang, Kuiyuan
Hua, Zhongyun
Lan, Rushi
Zhang, Yushu
Guo, Yifang
contents Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause significant security challenges when maliciously misused. Therefore, there is an urgent need to detect these synthetic speech signals. Phoneme features provide a powerful speech representation for deepfake detection. However, previous phoneme-based detection approaches typically focused on specific phonemes, overlooking temporal inconsistencies across the entire phoneme sequence. In this paper, we develop a new mechanism for detecting speech deepfakes by identifying the inconsistencies of phoneme-level speech features. We design an adaptive phoneme pooling technique that extracts sample-specific phoneme-level features from frame-level speech data. By applying this technique to features extracted by pre-trained audio models on previously unseen deepfake datasets, we demonstrate that deepfake samples often exhibit phoneme-level inconsistencies when compared to genuine speech. To further enhance detection accuracy, we propose a deepfake detector that uses a graph attention network to model the temporal dependencies of phoneme-level features. Additionally, we introduce a random phoneme substitution augmentation technique to increase feature diversity during training. Extensive experiments on four benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art detection methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12619
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
Zhang, Kuiyuan
Hua, Zhongyun
Lan, Rushi
Zhang, Yushu
Guo, Yifang
Sound
Artificial Intelligence
Audio and Speech Processing
Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause significant security challenges when maliciously misused. Therefore, there is an urgent need to detect these synthetic speech signals. Phoneme features provide a powerful speech representation for deepfake detection. However, previous phoneme-based detection approaches typically focused on specific phonemes, overlooking temporal inconsistencies across the entire phoneme sequence. In this paper, we develop a new mechanism for detecting speech deepfakes by identifying the inconsistencies of phoneme-level speech features. We design an adaptive phoneme pooling technique that extracts sample-specific phoneme-level features from frame-level speech data. By applying this technique to features extracted by pre-trained audio models on previously unseen deepfake datasets, we demonstrate that deepfake samples often exhibit phoneme-level inconsistencies when compared to genuine speech. To further enhance detection accuracy, we propose a deepfake detector that uses a graph attention network to model the temporal dependencies of phoneme-level features. Additionally, we introduce a random phoneme substitution augmentation technique to increase feature diversity during training. Extensive experiments on four benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art detection methods.
title Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.12619