Saved in:
Bibliographic Details
Main Authors: Chen, Muxi, Liu, Yi, Yi, Jian, Xu, Changran, Lai, Qiuxia, Wang, Hongliang, Ho, Tsung-Yi, Xu, Qiang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.05125
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929561593184256
author Chen, Muxi
Liu, Yi
Yi, Jian
Xu, Changran
Lai, Qiuxia
Wang, Hongliang
Ho, Tsung-Yi
Xu, Qiang
author_facet Chen, Muxi
Liu, Yi
Yi, Jian
Xu, Changran
Lai, Qiuxia
Wang, Hongliang
Ho, Tsung-Yi
Xu, Qiang
contents In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first, focusing on image qualities such as aesthetics and realism, and second, examining text conditions through concept coverage and fairness. We introduce an innovative aesthetic score prediction model that assesses the visual appeal of generated images and unveils the first dataset marked with low-quality regions in generated human images to facilitate automatic defect detection. Our exploration into concept coverage probes the model's effectiveness in interpreting and rendering text-based concepts accurately, while our analysis of fairness reveals biases in model outputs, with an emphasis on gender, race, and age. While our study is grounded in human imagery, this dual-faceted approach is designed with the flexibility to be applicable to other forms of image generation, enhancing our understanding of generative models and paving the way to the next generation of more sophisticated, contextually aware, and ethically attuned generative models. Code and data, including the dataset annotated with defective areas, are available at \href{https://github.com/cure-lab/EvaluateAIGC}{https://github.com/cure-lab/EvaluateAIGC}.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis
Chen, Muxi
Liu, Yi
Yi, Jian
Xu, Changran
Lai, Qiuxia
Wang, Hongliang
Ho, Tsung-Yi
Xu, Qiang
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first, focusing on image qualities such as aesthetics and realism, and second, examining text conditions through concept coverage and fairness. We introduce an innovative aesthetic score prediction model that assesses the visual appeal of generated images and unveils the first dataset marked with low-quality regions in generated human images to facilitate automatic defect detection. Our exploration into concept coverage probes the model's effectiveness in interpreting and rendering text-based concepts accurately, while our analysis of fairness reveals biases in model outputs, with an emphasis on gender, race, and age. While our study is grounded in human imagery, this dual-faceted approach is designed with the flexibility to be applicable to other forms of image generation, enhancing our understanding of generative models and paving the way to the next generation of more sophisticated, contextually aware, and ethically attuned generative models. Code and data, including the dataset annotated with defective areas, are available at \href{https://github.com/cure-lab/EvaluateAIGC}{https://github.com/cure-lab/EvaluateAIGC}.
title Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2403.05125