Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Taehoon, Choi, Jongwook, Jeong, Yonghyun, Noh, Haeun, Yoo, Jaejun, Baek, Seungryul, Choi, Jongwon
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918088141701120
author Kim, Taehoon
Choi, Jongwook
Jeong, Yonghyun
Noh, Haeun
Yoo, Jaejun
Baek, Seungryul
Choi, Jongwon
author_facet Kim, Taehoon
Choi, Jongwook
Jeong, Yonghyun
Noh, Haeun
Yoo, Jaejun
Baek, Seungryul
Choi, Jongwon
contents We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. Traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in the failure to detect temporal artifacts in the pixel plane. Our approach performs a 1D Fourier transform on the time axis for each pixel, extracting features highly sensitive to temporal inconsistencies, especially in areas prone to unnatural movements. To precisely locate regions containing the temporal artifacts, we introduce an attention proposal module trained in an end-to-end manner. Additionally, our joint transformer module effectively integrates pixel-wise temporal frequency features with spatio-temporal context features, expanding the range of detectable forgery artifacts. Our framework represents a significant advancement in deepfake video detection, providing robust performance across diverse and challenging detection scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
Kim, Taehoon
Choi, Jongwook
Jeong, Yonghyun
Noh, Haeun
Yoo, Jaejun
Baek, Seungryul
Choi, Jongwon
Computer Vision and Pattern Recognition
Artificial Intelligence
We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. Traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in the failure to detect temporal artifacts in the pixel plane. Our approach performs a 1D Fourier transform on the time axis for each pixel, extracting features highly sensitive to temporal inconsistencies, especially in areas prone to unnatural movements. To precisely locate regions containing the temporal artifacts, we introduce an attention proposal module trained in an end-to-end manner. Additionally, our joint transformer module effectively integrates pixel-wise temporal frequency features with spatio-temporal context features, expanding the range of detectable forgery artifacts. Our framework represents a significant advancement in deepfake video detection, providing robust performance across diverse and challenging detection scenarios.
title Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.02398