Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cobo, Alejandro, Valle, Roberto, Buenaposada, José Miguel, Baumela, Luis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915928730501120
author Cobo, Alejandro
Valle, Roberto
Buenaposada, José Miguel
Baumela, Luis
author_facet Cobo, Alejandro
Valle, Roberto
Buenaposada, José Miguel
Baumela, Luis
contents Generalizing deepfake detection to unseen manipulations remains a key challenge. A recent approach to tackle this issue is to train a network with pristine face images that have been manipulated with hand-crafted artifacts to extract more generalizable clues. While effective for static images, extending this to the video domain is an open issue. Existing methods model temporal artifacts as frame-to-frame instabilities, overlooking a key vulnerability: the violation of natural motion dependencies between different facial regions. In this paper, we propose a synthetic video generation method that creates training data with subtle kinematic inconsistencies. We train an autoencoder to decompose facial landmark configurations into motion bases. By manipulating these bases, we selectively break the natural correlations in facial movements and introduce these artifacts into pristine videos via face morphing. A network trained on our data learns to spot these sophisticated biomechanical flaws, achieving state-of-the-art generalization results on several popular benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04175
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection
Cobo, Alejandro
Valle, Roberto
Buenaposada, José Miguel
Baumela, Luis
Computer Vision and Pattern Recognition
Generalizing deepfake detection to unseen manipulations remains a key challenge. A recent approach to tackle this issue is to train a network with pristine face images that have been manipulated with hand-crafted artifacts to extract more generalizable clues. While effective for static images, extending this to the video domain is an open issue. Existing methods model temporal artifacts as frame-to-frame instabilities, overlooking a key vulnerability: the violation of natural motion dependencies between different facial regions. In this paper, we propose a synthetic video generation method that creates training data with subtle kinematic inconsistencies. We train an autoencoder to decompose facial landmark configurations into motion bases. By manipulating these bases, we selectively break the natural correlations in facial movements and introduce these artifacts into pristine videos via face morphing. A network trained on our data learns to spot these sophisticated biomechanical flaws, achieving state-of-the-art generalization results on several popular benchmarks.
title Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04175