Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wenhan, Karaoglu, Sezer, Gevers, Theo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914202059276288
author Chen, Wenhan
Karaoglu, Sezer
Gevers, Theo
author_facet Chen, Wenhan
Karaoglu, Sezer
Gevers, Theo
contents Recent advances in diffusion-based generation techniques enable AI models to produce highly realistic videos, heightening the need for reliable detection mechanisms. However, existing detection methods provide only limited exploration of the 3D geometric patterns present in generated videos. In this paper, we use vanishing points as an explicit representation of 3D geometry patterns, revealing fundamental discrepancies in geometric consistency between real and AI-generated videos. We introduce Grab-3D, a geometry-aware transformer framework for detecting AI-generated videos based on 3D geometric temporal consistency. To enable reliable evaluation, we construct an AI-generated video dataset of static scenes, allowing stable 3D geometric feature extraction. We propose a geometry-aware transformer equipped with geometric positional encoding, temporal-geometric attention, and an EMA-based geometric classifier head to explicitly inject 3D geometric awareness into temporal modeling. Experiments demonstrate that Grab-3D significantly outperforms state-of-the-art detectors, achieving robust cross-domain generalization to unseen generators.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13665
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
Chen, Wenhan
Karaoglu, Sezer
Gevers, Theo
Computer Vision and Pattern Recognition
Recent advances in diffusion-based generation techniques enable AI models to produce highly realistic videos, heightening the need for reliable detection mechanisms. However, existing detection methods provide only limited exploration of the 3D geometric patterns present in generated videos. In this paper, we use vanishing points as an explicit representation of 3D geometry patterns, revealing fundamental discrepancies in geometric consistency between real and AI-generated videos. We introduce Grab-3D, a geometry-aware transformer framework for detecting AI-generated videos based on 3D geometric temporal consistency. To enable reliable evaluation, we construct an AI-generated video dataset of static scenes, allowing stable 3D geometric feature extraction. We propose a geometry-aware transformer equipped with geometric positional encoding, temporal-geometric attention, and an EMA-based geometric classifier head to explicitly inject 3D geometric awareness into temporal modeling. Experiments demonstrate that Grab-3D significantly outperforms state-of-the-art detectors, achieving robust cross-domain generalization to unseen generators.
title Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13665