Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Jaemin, Hu, Yushi, Garg, Roopal, Anderson, Peter, Krishna, Ranjay, Baldridge, Jason, Bansal, Mohit, Pont-Tuset, Jordi, Wang, Su
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911797083111424
author Cho, Jaemin
Hu, Yushi
Garg, Roopal
Anderson, Peter
Krishna, Ranjay
Baldridge, Jason
Bansal, Mohit
Pont-Tuset, Jordi
Wang, Su
author_facet Cho, Jaemin
Hu, Yushi
Garg, Roopal
Anderson, Peter
Krishna, Ranjay
Baldridge, Jason
Bansal, Mohit
Pont-Tuset, Jordi
Wang, Su
contents Evaluating text-to-image models is notoriously difficult. A strong recent approach for assessing text-image faithfulness is based on QG/A (question generation and answering), which uses pre-trained foundational models to automatically generate a set of questions and answers from the prompt, and output images are scored based on whether these answers extracted with a visual question answering model are consistent with the prompt-based answers. This kind of evaluation is naturally dependent on the quality of the underlying QG and VQA models. We identify and address several reliability challenges in existing QG/A work: (a) QG questions should respect the prompt (avoiding hallucinations, duplications, and omissions) and (b) VQA answers should be consistent (not asserting that there is no motorcycle in an image while also claiming the motorcycle is blue). We address these issues with Davidsonian Scene Graph (DSG), an empirically grounded evaluation framework inspired by formal semantics, which is adaptable to any QG/A frameworks. DSG produces atomic and unique questions organized in dependency graphs, which (i) ensure appropriate semantic coverage and (ii) sidestep inconsistent answers. With extensive experimentation and human evaluation on a range of model configurations (LLM, VQA, and T2I), we empirically demonstrate that DSG addresses the challenges noted above. Finally, we present DSG-1k, an open-sourced evaluation benchmark that includes 1,060 prompts, covering a wide range of fine-grained semantic categories with a balanced distribution. We release the DSG-1k prompts and the corresponding DSG questions.
format Preprint
id arxiv_https___arxiv_org_abs_2310_18235
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
Cho, Jaemin
Hu, Yushi
Garg, Roopal
Anderson, Peter
Krishna, Ranjay
Baldridge, Jason
Bansal, Mohit
Pont-Tuset, Jordi
Wang, Su
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Evaluating text-to-image models is notoriously difficult. A strong recent approach for assessing text-image faithfulness is based on QG/A (question generation and answering), which uses pre-trained foundational models to automatically generate a set of questions and answers from the prompt, and output images are scored based on whether these answers extracted with a visual question answering model are consistent with the prompt-based answers. This kind of evaluation is naturally dependent on the quality of the underlying QG and VQA models. We identify and address several reliability challenges in existing QG/A work: (a) QG questions should respect the prompt (avoiding hallucinations, duplications, and omissions) and (b) VQA answers should be consistent (not asserting that there is no motorcycle in an image while also claiming the motorcycle is blue). We address these issues with Davidsonian Scene Graph (DSG), an empirically grounded evaluation framework inspired by formal semantics, which is adaptable to any QG/A frameworks. DSG produces atomic and unique questions organized in dependency graphs, which (i) ensure appropriate semantic coverage and (ii) sidestep inconsistent answers. With extensive experimentation and human evaluation on a range of model configurations (LLM, VQA, and T2I), we empirically demonstrate that DSG addresses the challenges noted above. Finally, we present DSG-1k, an open-sourced evaluation benchmark that includes 1,060 prompts, covering a wide range of fine-grained semantic categories with a balanced distribution. We release the DSG-1k prompts and the corresponding DSG questions.
title Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2310.18235