Time Step Generating: A Universal Synthesized Deepfake Image Detector

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Ziyue, Liu, Haoyuan, Peng, Dingjie, Jing, Luoxu, Watanabe, Hiroshi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913581430210560
author Zeng, Ziyue
Liu, Haoyuan
Peng, Dingjie
Jing, Luoxu
Watanabe, Hiroshi
author_facet Zeng, Ziyue
Liu, Haoyuan
Peng, Dingjie
Jing, Luoxu
Watanabe, Hiroshi
contents Currently, high-fidelity text-to-image models are developed in an accelerating pace. Among them, Diffusion Models have led to a remarkable improvement in the quality of image generation, making it vary challenging to distinguish between real and synthesized images. It simultaneously raises serious concerns regarding privacy and security. Some methods are proposed to distinguish the diffusion model generated images through reconstructing. However, the inversion and denoising processes are time-consuming and heavily reliant on the pre-trained generative model. Consequently, if the pre-trained generative model meet the problem of out-of-domain, the detection performance declines. To address this issue, we propose a universal synthetic image detector Time Step Generating (TSG), which does not rely on pre-trained models' reconstructing ability, specific datasets, or sampling algorithms. Our method utilizes a pre-trained diffusion model's network as a feature extractor to capture fine-grained details, focusing on the subtle differences between real and synthetic images. By controlling the time step t of the network input, we can effectively extract these distinguishing detail features. Then, those features can be passed through a classifier (i.e. Resnet), which efficiently detects whether an image is synthetic or real. We test the proposed TSG on the large-scale GenImage benchmark and it achieves significant improvements in both accuracy and generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11016
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Time Step Generating: A Universal Synthesized Deepfake Image Detector
Zeng, Ziyue
Liu, Haoyuan
Peng, Dingjie
Jing, Luoxu
Watanabe, Hiroshi
Computer Vision and Pattern Recognition
Artificial Intelligence
62H30, 68T07
I.4.9; I.4.7; I.5.2
Currently, high-fidelity text-to-image models are developed in an accelerating pace. Among them, Diffusion Models have led to a remarkable improvement in the quality of image generation, making it vary challenging to distinguish between real and synthesized images. It simultaneously raises serious concerns regarding privacy and security. Some methods are proposed to distinguish the diffusion model generated images through reconstructing. However, the inversion and denoising processes are time-consuming and heavily reliant on the pre-trained generative model. Consequently, if the pre-trained generative model meet the problem of out-of-domain, the detection performance declines. To address this issue, we propose a universal synthetic image detector Time Step Generating (TSG), which does not rely on pre-trained models' reconstructing ability, specific datasets, or sampling algorithms. Our method utilizes a pre-trained diffusion model's network as a feature extractor to capture fine-grained details, focusing on the subtle differences between real and synthetic images. By controlling the time step t of the network input, we can effectively extract these distinguishing detail features. Then, those features can be passed through a classifier (i.e. Resnet), which efficiently detects whether an image is synthetic or real. We test the proposed TSG on the large-scale GenImage benchmark and it achieves significant improvements in both accuracy and generalizability.
title Time Step Generating: A Universal Synthesized Deepfake Image Detector
topic Computer Vision and Pattern Recognition
Artificial Intelligence
62H30, 68T07
I.4.9; I.4.7; I.5.2
url https://arxiv.org/abs/2411.11016