PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Datta, Soumyya Kanti, Ranga, Tanvi, Sun, Chengzhe, Lyu, Siwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911214062272512
author Datta, Soumyya Kanti
Ranga, Tanvi
Sun, Chengzhe
Lyu, Siwei
author_facet Datta, Soumyya Kanti
Ranga, Tanvi
Sun, Chengzhe
Lyu, Siwei
contents The rise of manipulated media has made deepfakes a particularly insidious threat, involving various generative manipulations such as lip-sync modifications, face-swaps, and avatar-driven facial synthesis. Conventional detection methods, which predominantly depend on manually designed phoneme-viseme alignment thresholds, fundamental frame-level consistency checks, or a unimodal detection strategy, inadequately identify modern-day deepfakes generated by advanced generative models such as GANs, diffusion models, and neural rendering techniques. These advanced techniques generate nearly perfect individual frames yet inadvertently create minor temporal discrepancies frequently overlooked by traditional detectors. We present a novel multimodal audio-visual framework, Phoneme-Temporal and Identity-Dynamic Analysis(PIA), incorporating language, dynamic face motion, and facial identification cues to address these limitations. We utilize phoneme sequences, lip geometry data, and advanced facial identity embeddings. This integrated method significantly improves the detection of subtle deepfake alterations by identifying inconsistencies across multiple complementary modalities. Code is available at https://github.com/skrantidatta/PIA
format Preprint
id arxiv_https___arxiv_org_abs_2510_14241
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
Datta, Soumyya Kanti
Ranga, Tanvi
Sun, Chengzhe
Lyu, Siwei
Computer Vision and Pattern Recognition
The rise of manipulated media has made deepfakes a particularly insidious threat, involving various generative manipulations such as lip-sync modifications, face-swaps, and avatar-driven facial synthesis. Conventional detection methods, which predominantly depend on manually designed phoneme-viseme alignment thresholds, fundamental frame-level consistency checks, or a unimodal detection strategy, inadequately identify modern-day deepfakes generated by advanced generative models such as GANs, diffusion models, and neural rendering techniques. These advanced techniques generate nearly perfect individual frames yet inadvertently create minor temporal discrepancies frequently overlooked by traditional detectors. We present a novel multimodal audio-visual framework, Phoneme-Temporal and Identity-Dynamic Analysis(PIA), incorporating language, dynamic face motion, and facial identification cues to address these limitations. We utilize phoneme sequences, lip geometry data, and advanced facial identity embeddings. This integrated method significantly improves the detection of subtle deepfake alterations by identifying inconsistencies across multiple complementary modalities. Code is available at https://github.com/skrantidatta/PIA
title PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14241