Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wiles, Olivia, Zhang, Chuhan, Albuquerque, Isabela, Kajić, Ivana, Wang, Su, Bugliarello, Emanuele, Onoe, Yasumasa, Papalampidi, Pinelopi, Ktena, Ira, Knutsen, Chris, Rashtchian, Cyrus, Nawalgaria, Anant, Pont-Tuset, Jordi, Nematzadeh, Aida
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916654094483456
author Wiles, Olivia
Zhang, Chuhan
Albuquerque, Isabela
Kajić, Ivana
Wang, Su
Bugliarello, Emanuele
Onoe, Yasumasa
Papalampidi, Pinelopi
Ktena, Ira
Knutsen, Chris
Rashtchian, Cyrus
Nawalgaria, Anant
Pont-Tuset, Jordi
Nematzadeh, Aida
author_facet Wiles, Olivia
Zhang, Chuhan
Albuquerque, Isabela
Kajić, Ivana
Wang, Su
Bugliarello, Emanuele
Onoe, Yasumasa
Papalampidi, Pinelopi
Ktena, Ira
Knutsen, Chris
Rashtchian, Cyrus
Nawalgaria, Anant
Pont-Tuset, Jordi
Nematzadeh, Aida
contents While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I alignment by proposing metrics, benchmarks, and templates for collecting human judgements, the quality of these components is not systematically measured. Human-rated prompt sets are generally small and the reliability of the ratings -- and thereby the prompt set used to compare models -- is not evaluated. We address this gap by performing an extensive study evaluating auto-eval metrics and human templates. We provide three main contributions: (1) We introduce a comprehensive skills-based benchmark that can discriminate models across different human templates. This skills-based benchmark categorises prompts into sub-skills, allowing a practitioner to pinpoint not only which skills are challenging, but at what level of complexity a skill becomes challenging. (2) We gather human ratings across four templates and four T2I models for a total of >100K annotations. This allows us to understand where differences arise due to inherent ambiguity in the prompt and where they arise due to differences in metric and model quality. (3) Finally, we introduce a new QA-based auto-eval metric that is better correlated with human ratings than existing metrics for our new dataset, across different human templates, and on TIFA160.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16820
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings
Wiles, Olivia
Zhang, Chuhan
Albuquerque, Isabela
Kajić, Ivana
Wang, Su
Bugliarello, Emanuele
Onoe, Yasumasa
Papalampidi, Pinelopi
Ktena, Ira
Knutsen, Chris
Rashtchian, Cyrus
Nawalgaria, Anant
Pont-Tuset, Jordi
Nematzadeh, Aida
Computer Vision and Pattern Recognition
While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I alignment by proposing metrics, benchmarks, and templates for collecting human judgements, the quality of these components is not systematically measured. Human-rated prompt sets are generally small and the reliability of the ratings -- and thereby the prompt set used to compare models -- is not evaluated. We address this gap by performing an extensive study evaluating auto-eval metrics and human templates. We provide three main contributions: (1) We introduce a comprehensive skills-based benchmark that can discriminate models across different human templates. This skills-based benchmark categorises prompts into sub-skills, allowing a practitioner to pinpoint not only which skills are challenging, but at what level of complexity a skill becomes challenging. (2) We gather human ratings across four templates and four T2I models for a total of >100K annotations. This allows us to understand where differences arise due to inherent ambiguity in the prompt and where they arise due to differences in metric and model quality. (3) Finally, we introduce a new QA-based auto-eval metric that is better correlated with human ratings than existing metrics for our new dataset, across different human templates, and on TIFA160.
title Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.16820