SFANet: Spatial-Frequency Attention Network for Deepfake Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ahire, Vrushank, Muley, Aniruddh, Zample, Shivam, Verma, Siddharth, Menon, Pranav, Madan, Surbhi, Dhall, Abhinav
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912630677962752
author Ahire, Vrushank
Muley, Aniruddh
Zample, Shivam
Verma, Siddharth
Menon, Pranav
Madan, Surbhi
Dhall, Abhinav
author_facet Ahire, Vrushank
Muley, Aniruddh
Zample, Shivam
Verma, Siddharth
Menon, Pranav
Madan, Surbhi
Dhall, Abhinav
contents Detecting manipulated media has now become a pressing issue with the recent rise of deepfakes. Most existing approaches fail to generalize across diverse datasets and generation techniques. We thus propose a novel ensemble framework, combining the strengths of transformer-based architectures, such as Swin Transformers and ViTs, and texture-based methods, to achieve better detection accuracy and robustness. Our method introduces innovative data-splitting, sequential training, frequency splitting, patch-based attention, and face segmentation techniques to handle dataset imbalances, enhance high-impact regions (e.g., eyes and mouth), and improve generalization. Our model achieves state-of-the-art performance when tested on the DFWild-Cup dataset, a diverse subset of eight deepfake datasets. The ensemble benefits from the complementarity of these approaches, with transformers excelling in global feature extraction and texturebased methods providing interpretability. This work demonstrates that hybrid models can effectively address the evolving challenges of deepfake detection, offering a robust solution for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SFANet: Spatial-Frequency Attention Network for Deepfake Detection
Ahire, Vrushank
Muley, Aniruddh
Zample, Shivam
Verma, Siddharth
Menon, Pranav
Madan, Surbhi
Dhall, Abhinav
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Detecting manipulated media has now become a pressing issue with the recent rise of deepfakes. Most existing approaches fail to generalize across diverse datasets and generation techniques. We thus propose a novel ensemble framework, combining the strengths of transformer-based architectures, such as Swin Transformers and ViTs, and texture-based methods, to achieve better detection accuracy and robustness. Our method introduces innovative data-splitting, sequential training, frequency splitting, patch-based attention, and face segmentation techniques to handle dataset imbalances, enhance high-impact regions (e.g., eyes and mouth), and improve generalization. Our model achieves state-of-the-art performance when tested on the DFWild-Cup dataset, a diverse subset of eight deepfake datasets. The ensemble benefits from the complementarity of these approaches, with transformers excelling in global feature extraction and texturebased methods providing interpretability. This work demonstrates that hybrid models can effectively address the evolving challenges of deepfake detection, offering a robust solution for real-world applications.
title SFANet: Spatial-Frequency Attention Network for Deepfake Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2510.04630