Reduced Spatial Dependency for More General Video-level Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Beilin, Xu, Xuan, Zhang, Yufei, You, Weike, Zhou, Linna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910859496783872
author Chu, Beilin
Xu, Xuan
Zhang, Yufei
You, Weike
Zhou, Linna
author_facet Chu, Beilin
Xu, Xuan
Zhang, Yufei
You, Weike
Zhou, Linna
contents As one of the prominent AI-generated content, Deepfake has raised significant safety concerns. Although it has been demonstrated that temporal consistency cues offer better generalization capability, existing methods based on CNNs inevitably introduce spatial bias, which hinders the extraction of intrinsic temporal features. To address this issue, we propose a novel method called Spatial Dependency Reduction (SDR), which integrates common temporal consistency features from multiple spatially-perturbed clusters, to reduce the dependency of the model on spatial information. Specifically, we design multiple Spatial Perturbation Branch (SPB) to construct spatially-perturbed feature clusters. Subsequently, we utilize the theory of mutual information and propose a Task-Relevant Feature Integration (TRFI) module to capture temporal features residing in similar latent space from these clusters. Finally, the integrated feature is fed into a temporal transformer to capture long-range dependencies. Extensive benchmarks and ablation studies demonstrate the effectiveness and rationale of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reduced Spatial Dependency for More General Video-level Deepfake Detection
Chu, Beilin
Xu, Xuan
Zhang, Yufei
You, Weike
Zhou, Linna
Computer Vision and Pattern Recognition
Cryptography and Security
As one of the prominent AI-generated content, Deepfake has raised significant safety concerns. Although it has been demonstrated that temporal consistency cues offer better generalization capability, existing methods based on CNNs inevitably introduce spatial bias, which hinders the extraction of intrinsic temporal features. To address this issue, we propose a novel method called Spatial Dependency Reduction (SDR), which integrates common temporal consistency features from multiple spatially-perturbed clusters, to reduce the dependency of the model on spatial information. Specifically, we design multiple Spatial Perturbation Branch (SPB) to construct spatially-perturbed feature clusters. Subsequently, we utilize the theory of mutual information and propose a Task-Relevant Feature Integration (TRFI) module to capture temporal features residing in similar latent space from these clusters. Finally, the integrated feature is fed into a temporal transformer to capture long-range dependencies. Extensive benchmarks and ablation studies demonstrate the effectiveness and rationale of our approach.
title Reduced Spatial Dependency for More General Video-level Deepfake Detection
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2503.03270