Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918088141701120 |
|---|---|
| author | Kim, Taehoon Choi, Jongwook Jeong, Yonghyun Noh, Haeun Yoo, Jaejun Baek, Seungryul Choi, Jongwon |
| author_facet | Kim, Taehoon Choi, Jongwook Jeong, Yonghyun Noh, Haeun Yoo, Jaejun Baek, Seungryul Choi, Jongwon |
| contents | We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. Traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in the failure to detect temporal artifacts in the pixel plane. Our approach performs a 1D Fourier transform on the time axis for each pixel, extracting features highly sensitive to temporal inconsistencies, especially in areas prone to unnatural movements. To precisely locate regions containing the temporal artifacts, we introduce an attention proposal module trained in an end-to-end manner. Additionally, our joint transformer module effectively integrates pixel-wise temporal frequency features with spatio-temporal context features, expanding the range of detectable forgery artifacts. Our framework represents a significant advancement in deepfake video detection, providing robust performance across diverse and challenging detection scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_02398 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection Kim, Taehoon Choi, Jongwook Jeong, Yonghyun Noh, Haeun Yoo, Jaejun Baek, Seungryul Choi, Jongwon Computer Vision and Pattern Recognition Artificial Intelligence We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. Traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in the failure to detect temporal artifacts in the pixel plane. Our approach performs a 1D Fourier transform on the time axis for each pixel, extracting features highly sensitive to temporal inconsistencies, especially in areas prone to unnatural movements. To precisely locate regions containing the temporal artifacts, we introduce an attention proposal module trained in an end-to-end manner. Additionally, our joint transformer module effectively integrates pixel-wise temporal frequency features with spatio-temporal context features, expanding the range of detectable forgery artifacts. Our framework represents a significant advancement in deepfake video detection, providing robust performance across diverse and challenging detection scenarios. |
| title | Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2507.02398 |