Saved in:
Bibliographic Details
Main Authors: Chen, Hao, Xu, Junnan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.16808
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911716548280320
author Chen, Hao
Xu, Junnan
author_facet Chen, Hao
Xu, Junnan
contents Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different approach. Authentic lip motion is constrained by tissue mechanics and neuromuscular bandwidth; current generators typically do not impose these constraints, producing trajectories with elevated variance in velocity, acceleration, and jerk that real speech does not exhibit. We exploit this signal, which we term temporal lip jitter, by computing kinematic statistics from 64 perioral landmarks over short sliding windows and feeding them into a lightweight three-branch network. The model uses only landmark coordinates: no pixels, no audio, and no voiceprint data. We train only on English data and test in a zero-shot setting on five unseen generators and seven languages.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16808
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling
Chen, Hao
Xu, Junnan
Computer Vision and Pattern Recognition
Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different approach. Authentic lip motion is constrained by tissue mechanics and neuromuscular bandwidth; current generators typically do not impose these constraints, producing trajectories with elevated variance in velocity, acceleration, and jerk that real speech does not exhibit. We exploit this signal, which we term temporal lip jitter, by computing kinematic statistics from 64 perioral landmarks over short sliding windows and feeding them into a lightweight three-branch network. The model uses only landmark coordinates: no pixels, no audio, and no voiceprint data. We train only on English data and test in a zero-shot setting on five unseen generators and seven languages.
title BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.16808