Explainable Synthetic Image Detection through Diffusion Timestep Ensembling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yixin, Zhang, Feiran, Shi, Tianyuan, Yin, Ruicheng, Wang, Zhenghua, Gan, Zhenliang, Wang, Xiaohua, Lv, Changze, Zheng, Xiaoqing, Huang, Xuanjing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909707664359424
author Wu, Yixin
Zhang, Feiran
Shi, Tianyuan
Yin, Ruicheng
Wang, Zhenghua
Gan, Zhenliang
Wang, Xiaohua
Lv, Changze
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Wu, Yixin
Zhang, Feiran
Shi, Tianyuan
Yin, Ruicheng
Wang, Zhenghua
Gan, Zhenliang
Wang, Xiaohua
Lv, Changze
Zheng, Xiaoqing
Huang, Xuanjing
contents Recent advances in diffusion models have enabled the creation of deceptively real images, posing significant security risks when misused. In this study, we empirically show that different timesteps of DDIM inversion reveal varying subtle distinctions between synthetic and real images that are extractable for detection, in the forms of such as Fourier power spectrum high-frequency discrepancies and inter-pixel variance distributions. Based on these observations, we propose a novel synthetic image detection method that directly utilizes features of intermediately noised images by training an ensemble on multiple noised timesteps, circumventing conventional reconstruction-based strategies. To enhance human comprehension, we introduce a metric-grounded explanation generation and refinement module to identify and explain AI-generated flaws. Additionally, we construct the GenHard and GenExplain benchmarks to provide detection samples of greater difficulty and high-quality rationales for fake images. Extensive experiments show that our method achieves state-of-the-art performance with 98.91% and 95.89% detection accuracy on regular and challenging samples respectively, and demonstrates generalizability and robustness. Our code and datasets are available at https://github.com/Shadowlized/ESIDE.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06201
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explainable Synthetic Image Detection through Diffusion Timestep Ensembling
Wu, Yixin
Zhang, Feiran
Shi, Tianyuan
Yin, Ruicheng
Wang, Zhenghua
Gan, Zhenliang
Wang, Xiaohua
Lv, Changze
Zheng, Xiaoqing
Huang, Xuanjing
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent advances in diffusion models have enabled the creation of deceptively real images, posing significant security risks when misused. In this study, we empirically show that different timesteps of DDIM inversion reveal varying subtle distinctions between synthetic and real images that are extractable for detection, in the forms of such as Fourier power spectrum high-frequency discrepancies and inter-pixel variance distributions. Based on these observations, we propose a novel synthetic image detection method that directly utilizes features of intermediately noised images by training an ensemble on multiple noised timesteps, circumventing conventional reconstruction-based strategies. To enhance human comprehension, we introduce a metric-grounded explanation generation and refinement module to identify and explain AI-generated flaws. Additionally, we construct the GenHard and GenExplain benchmarks to provide detection samples of greater difficulty and high-quality rationales for fake images. Extensive experiments show that our method achieves state-of-the-art performance with 98.91% and 95.89% detection accuracy on regular and challenging samples respectively, and demonstrates generalizability and robustness. Our code and datasets are available at https://github.com/Shadowlized/ESIDE.
title Explainable Synthetic Image Detection through Diffusion Timestep Ensembling
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.06201