Deepfake Detection in Social Media: A Temporal Artifact Analysis Using 3D Convolutional Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rashidi, Mohammadreza, Ali, Raja Hashim, Rahman, Sami Ur
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916021073346560
author Rashidi, Mohammadreza
Ali, Raja Hashim
Rahman, Sami Ur
author_facet Rashidi, Mohammadreza
Ali, Raja Hashim
Rahman, Sami Ur
contents Synthetic facial videos have proliferated across social media faster than platform moderation can respond, raising the cost of disinformation and identity-based attacks. Frame-level deepfake detectors degrade sharply as generator quality increases; high-quality 128x128 GAN output cuts spatial-only accuracy by five percentage points while leaving temporal inconsistencies largely intact. We address this gap with a 3D Convolutional Neural Network detector based on R3D-18, trained with a composite loss that combines binary cross-entropy with a temporal-consistency regularizer. The model processes 16-frame clips from the DeepfakeTIMIT dataset and is initialized from Kinetics-400 action-recognition weights. We report 92.8% accuracy on intra-dataset evaluation at 128x128 resolution; cross-dataset transfer to FaceForensics++ without fine-tuning reaches 76.4%, rising after minimal fine-tuning. Ablation studies show that transfer learning contributes 7.2 percentage points and face tracking adds 3.5 points, while temporal consistency regularization provides additional gains on high-quality fakes. The results establish that temporal artifacts generalize more broadly than spatial ones, providing a detection signal that survives social-media re-encoding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17573
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Deepfake Detection in Social Media: A Temporal Artifact Analysis Using 3D Convolutional Neural Networks
Rashidi, Mohammadreza
Ali, Raja Hashim
Rahman, Sami Ur
Computer Vision and Pattern Recognition
Cryptography and Security
I.4.9; I.2.10; H.3.1
Synthetic facial videos have proliferated across social media faster than platform moderation can respond, raising the cost of disinformation and identity-based attacks. Frame-level deepfake detectors degrade sharply as generator quality increases; high-quality 128x128 GAN output cuts spatial-only accuracy by five percentage points while leaving temporal inconsistencies largely intact. We address this gap with a 3D Convolutional Neural Network detector based on R3D-18, trained with a composite loss that combines binary cross-entropy with a temporal-consistency regularizer. The model processes 16-frame clips from the DeepfakeTIMIT dataset and is initialized from Kinetics-400 action-recognition weights. We report 92.8% accuracy on intra-dataset evaluation at 128x128 resolution; cross-dataset transfer to FaceForensics++ without fine-tuning reaches 76.4%, rising after minimal fine-tuning. Ablation studies show that transfer learning contributes 7.2 percentage points and face tracking adds 3.5 points, while temporal consistency regularization provides additional gains on high-quality fakes. The results establish that temporal artifacts generalize more broadly than spatial ones, providing a detection signal that survives social-media re-encoding.
title Deepfake Detection in Social Media: A Temporal Artifact Analysis Using 3D Convolutional Neural Networks
topic Computer Vision and Pattern Recognition
Cryptography and Security
I.4.9; I.2.10; H.3.1
url https://arxiv.org/abs/2605.17573