Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: You, Zebin, Zhang, Xinyu, Guo, Hanzhong, Wang, Jingdong, Li, Chongxuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913752407867392
author You, Zebin
Zhang, Xinyu
Guo, Hanzhong
Wang, Jingdong
Li, Chongxuan
author_facet You, Zebin
Zhang, Xinyu
Guo, Hanzhong
Wang, Jingdong
Li, Chongxuan
contents The ultimate goal of generative models is to perfectly capture the data distribution. For image generation, common metrics of visual quality (e.g., FID) and the perceived truthfulness of generated images seem to suggest that we are nearing this goal. However, through distribution classification tasks, we reveal that, from the perspective of neural network-based classifiers, even advanced diffusion models are still far from this goal. Specifically, classifiers are able to consistently and effortlessly distinguish real images from generated ones across various settings. Moreover, we uncover an intriguing discrepancy: classifiers can easily differentiate between diffusion models with comparable performance (e.g., U-ViT-H vs. DiT-XL), but struggle to distinguish between models within the same family but of different scales (e.g., EDM2-XS vs. EDM2-XXL). Our methodology carries several important implications. First, it naturally serves as a diagnostic tool for diffusion models by analyzing specific features of generated data. Second, it sheds light on the model autophagy disorder and offers insights into the use of generated data: augmenting real data with generated data is more effective than replacing it. Third, classifier guidance can significantly enhance the realism of generated images.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18029
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?
You, Zebin
Zhang, Xinyu
Guo, Hanzhong
Wang, Jingdong
Li, Chongxuan
Computer Vision and Pattern Recognition
Machine Learning
The ultimate goal of generative models is to perfectly capture the data distribution. For image generation, common metrics of visual quality (e.g., FID) and the perceived truthfulness of generated images seem to suggest that we are nearing this goal. However, through distribution classification tasks, we reveal that, from the perspective of neural network-based classifiers, even advanced diffusion models are still far from this goal. Specifically, classifiers are able to consistently and effortlessly distinguish real images from generated ones across various settings. Moreover, we uncover an intriguing discrepancy: classifiers can easily differentiate between diffusion models with comparable performance (e.g., U-ViT-H vs. DiT-XL), but struggle to distinguish between models within the same family but of different scales (e.g., EDM2-XS vs. EDM2-XXL). Our methodology carries several important implications. First, it naturally serves as a diagnostic tool for diffusion models by analyzing specific features of generated data. Second, it sheds light on the model autophagy disorder and offers insights into the use of generated data: augmenting real data with generated data is more effective than replacing it. Third, classifier guidance can significantly enhance the realism of generated images.
title Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.18029