Physics-Guided Deepfake Detection for Voice Authentication Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mohammadi, Alireza, Sood, Keshav, Thiruvady, Dhananjay, Nazari, Asef
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915657594961920
author Mohammadi, Alireza
Sood, Keshav
Thiruvady, Dhananjay
Nazari, Asef
author_facet Mohammadi, Alireza
Sood, Keshav
Thiruvady, Dhananjay
Nazari, Asef
contents Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling physics-guided deepfake detection with uncertainty-aware in edge learning. The framework fuses interpretable physics features modeling vocal tract dynamics with representations coming from a self-supervised learning module. The representations are then processed via a Multi-Modal Ensemble Architecture, followed by a Bayesian ensemble providing uncertainty estimates. Incorporating physics-based characteristics evaluations and uncertainty estimates of audio samples allows our proposed framework to remain robust to both advanced deepfake attacks and sophisticated control-plane poisoning, addressing the complete threat model for networked voice authentication.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Physics-Guided Deepfake Detection for Voice Authentication Systems
Mohammadi, Alireza
Sood, Keshav
Thiruvady, Dhananjay
Nazari, Asef
Sound
Artificial Intelligence
Audio and Speech Processing
Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling physics-guided deepfake detection with uncertainty-aware in edge learning. The framework fuses interpretable physics features modeling vocal tract dynamics with representations coming from a self-supervised learning module. The representations are then processed via a Multi-Modal Ensemble Architecture, followed by a Bayesian ensemble providing uncertainty estimates. Incorporating physics-based characteristics evaluations and uncertainty estimates of audio samples allows our proposed framework to remain robust to both advanced deepfake attacks and sophisticated control-plane poisoning, addressing the complete threat model for networked voice authentication.
title Physics-Guided Deepfake Detection for Voice Authentication Systems
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2512.06040