DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Saksham, Singh, Ashish, Thota, Srinivasarao, Singh, Sunil Kumar, Kumar, Chandan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915836400238592
author Kumar, Saksham
Singh, Ashish
Thota, Srinivasarao
Singh, Sunil Kumar
Kumar, Chandan
author_facet Kumar, Saksham
Singh, Ashish
Thota, Srinivasarao
Singh, Sunil Kumar
Kumar, Chandan
contents Deepfakes are major threats to the integrity of digital media. We propose DeiTFake, a DeiT-based transformer and a novel two-stage progressive training strategy with increasing augmentation complexity. The approach applies an initial transfer-learning phase with standard augmentations followed by a fine-tuning phase using advanced affine and deepfake-specific augmentations. DeiT's knowledge distillation model captures subtle manipulation artifacts, increasing robustness of the detection model. Trained on the OpenForensics dataset (190,335 images), DeiTFake achieves 98.71\% accuracy after stage one and 99.22\% accuracy with an AUROC of 0.9997, after stage two, outperforming the latest OpenForensics baselines. We analyze augmentation impact and training schedules, and provide practical benchmarks for facial deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training
Kumar, Saksham
Singh, Ashish
Thota, Srinivasarao
Singh, Sunil Kumar
Kumar, Chandan
Computer Vision and Pattern Recognition
Cryptography and Security
Deepfakes are major threats to the integrity of digital media. We propose DeiTFake, a DeiT-based transformer and a novel two-stage progressive training strategy with increasing augmentation complexity. The approach applies an initial transfer-learning phase with standard augmentations followed by a fine-tuning phase using advanced affine and deepfake-specific augmentations. DeiT's knowledge distillation model captures subtle manipulation artifacts, increasing robustness of the detection model. Trained on the OpenForensics dataset (190,335 images), DeiTFake achieves 98.71\% accuracy after stage one and 99.22\% accuracy with an AUROC of 0.9997, after stage two, outperforming the latest OpenForensics baselines. We analyze augmentation impact and training schedules, and provide practical benchmarks for facial deepfake detection.
title DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2511.12048