Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salvi, Davide, Negroni, Viola, Mandelli, Sara, Bestagini, Paolo, Tubaro, Stefano
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911051415552000
author Salvi, Davide
Negroni, Viola
Mandelli, Sara
Bestagini, Paolo
Tubaro, Stefano
author_facet Salvi, Davide
Negroni, Viola
Mandelli, Sara
Bestagini, Paolo
Tubaro, Stefano
contents Recent advances in generative AI have made the creation of speech deepfakes widely accessible, posing serious challenges to digital trust. To counter this, various speech deepfake detection strategies have been proposed, including Person-of-Interest (POI) approaches, which focus on identifying impersonations of specific individuals by modeling and analyzing their unique vocal traits. Despite their excellent performance, the existing methods offer limited granularity and lack interpretability. In this work, we propose a POI-based speech deepfake detection method that operates at the phoneme level. Our approach decomposes reference audio into phonemes to construct a detailed speaker profile. In inference, phonemes from a test sample are individually compared against this profile, enabling fine-grained detection of synthetic artifacts. The proposed method achieves comparable accuracy to traditional approaches while offering superior robustness and interpretability, key aspects in multimedia forensics. By focusing on phoneme analysis, this work explores a novel direction for explainable, speaker-centric deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
Salvi, Davide
Negroni, Viola
Mandelli, Sara
Bestagini, Paolo
Tubaro, Stefano
Sound
Audio and Speech Processing
Recent advances in generative AI have made the creation of speech deepfakes widely accessible, posing serious challenges to digital trust. To counter this, various speech deepfake detection strategies have been proposed, including Person-of-Interest (POI) approaches, which focus on identifying impersonations of specific individuals by modeling and analyzing their unique vocal traits. Despite their excellent performance, the existing methods offer limited granularity and lack interpretability. In this work, we propose a POI-based speech deepfake detection method that operates at the phoneme level. Our approach decomposes reference audio into phonemes to construct a detailed speaker profile. In inference, phonemes from a test sample are individually compared against this profile, enabling fine-grained detection of synthetic artifacts. The proposed method achieves comparable accuracy to traditional approaches while offering superior robustness and interpretability, key aspects in multimedia forensics. By focusing on phoneme analysis, this work explores a novel direction for explainable, speaker-centric deepfake detection.
title Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.08626