Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jang, Minsuk, Jeong, Hyunseo, Son, Minseok, Kim, Changick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912953363595264
author Jang, Minsuk
Jeong, Hyunseo
Son, Minseok
Kim, Changick
author_facet Jang, Minsuk
Jeong, Hyunseo
Son, Minseok
Kim, Changick
contents Context-based detection methods such as DetectGPT achieve strong generalization in identifying AI-generated text by evaluating content compatibility with a model's learned distribution. In contrast, existing image detectors rely on discriminative features from pretrained backbones such as CLIP, which implicitly capture generator-specific artifacts. However, as modern generative models rapidly advance in visual fidelity, the artifacts these detectors depend on are becoming increasingly subtle or absent, undermining their reliability. Masked AutoEncoders (MAE) are inherently trained to reconstruct masked patches from visible context, naturally modeling patch-level contextual plausibility akin to conditional probability estimation, while also serving as a powerful semantic feature extractor through its encoder. We propose CINEMAE, a novel architecture that exploits both capabilities of MAE for AI-generated image detection: we derive per-patch anomaly signals from the reconstruction mechanism and extract global semantic features from the encoder, fusing both context-based and feature-based cues for robust detection. CINEMAE achieves highly competitive mean accuracies of 96.63\% on GenImage and 93.96\% on AIGCDetectBenchmark, maintaining over 93\% accuracy even under JPEG compression at QF=50.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders
Jang, Minsuk
Jeong, Hyunseo
Son, Minseok
Kim, Changick
Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
68T07
I.4.8
Context-based detection methods such as DetectGPT achieve strong generalization in identifying AI-generated text by evaluating content compatibility with a model's learned distribution. In contrast, existing image detectors rely on discriminative features from pretrained backbones such as CLIP, which implicitly capture generator-specific artifacts. However, as modern generative models rapidly advance in visual fidelity, the artifacts these detectors depend on are becoming increasingly subtle or absent, undermining their reliability. Masked AutoEncoders (MAE) are inherently trained to reconstruct masked patches from visible context, naturally modeling patch-level contextual plausibility akin to conditional probability estimation, while also serving as a powerful semantic feature extractor through its encoder. We propose CINEMAE, a novel architecture that exploits both capabilities of MAE for AI-generated image detection: we derive per-patch anomaly signals from the reconstruction mechanism and extract global semantic features from the encoder, fusing both context-based and feature-based cues for robust detection. CINEMAE achieves highly competitive mean accuracies of 96.63\% on GenImage and 93.96\% on AIGCDetectBenchmark, maintaining over 93\% accuracy even under JPEG compression at QF=50.
title Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
68T07
I.4.8
url https://arxiv.org/abs/2511.06325