EvalGIM: A Library for Evaluating Generative Image Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hall, Melissa, Mañas, Oscar, Askari-Hemmat, Reyhane, Ibrahim, Mark, Ross, Candace, Astolfi, Pietro, Ifriqi, Tariq Berrada, Havasi, Marton, Benchetrit, Yohann, Ullrich, Karen, Braga, Carolina, Charnalia, Abhishek, Ryan, Maeve, Rabbat, Mike, Drozdzal, Michal, Verbeek, Jakob, Romero-Soriano, Adriana
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915069673078784
author Hall, Melissa
Mañas, Oscar
Askari-Hemmat, Reyhane
Ibrahim, Mark
Ross, Candace
Astolfi, Pietro
Ifriqi, Tariq Berrada
Havasi, Marton
Benchetrit, Yohann
Ullrich, Karen
Braga, Carolina
Charnalia, Abhishek
Ryan, Maeve
Rabbat, Mike
Drozdzal, Michal
Verbeek, Jakob
Romero-Soriano, Adriana
author_facet Hall, Melissa
Mañas, Oscar
Askari-Hemmat, Reyhane
Ibrahim, Mark
Ross, Candace
Astolfi, Pietro
Ifriqi, Tariq Berrada
Havasi, Marton
Benchetrit, Yohann
Ullrich, Karen
Braga, Carolina
Charnalia, Abhishek
Ryan, Maeve
Rabbat, Mike
Drozdzal, Michal
Verbeek, Jakob
Romero-Soriano, Adriana
contents As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound, there are few unified benchmarking libraries that provide a framework for performing evaluations across many datasets and metrics. Furthermore, the rapid introduction of increasingly robust benchmarking methods requires that evaluation libraries remain flexible to new datasets and metrics. Finally, there remains a gap in synthesizing evaluations in order to deliver actionable takeaways about model performance. To enable unified, flexible, and actionable evaluations, we introduce EvalGIM (pronounced ''EvalGym''), a library for evaluating generative image models. EvalGIM contains broad support for datasets and metrics used to measure quality, diversity, and consistency of text-to-image generative models. In addition, EvalGIM is designed with flexibility for user customization as a top priority and contains a structure that allows plug-and-play additions of new datasets and metrics. To enable actionable evaluation insights, we introduce ''Evaluation Exercises'' that highlight takeaways for specific evaluation questions. The Evaluation Exercises contain easy-to-use and reproducible implementations of two state-of-the-art evaluation methods of text-to-image generative models: consistency-diversity-realism Pareto Fronts and disaggregated measurements of performance disparities across groups. EvalGIM also contains Evaluation Exercises that introduce two new analysis methods for text-to-image generative models: robustness analyses of model rankings and balanced evaluations across different prompt styles. We encourage text-to-image model exploration with EvalGIM and invite contributions at https://github.com/facebookresearch/EvalGIM/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10604
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EvalGIM: A Library for Evaluating Generative Image Models
Hall, Melissa
Mañas, Oscar
Askari-Hemmat, Reyhane
Ibrahim, Mark
Ross, Candace
Astolfi, Pietro
Ifriqi, Tariq Berrada
Havasi, Marton
Benchetrit, Yohann
Ullrich, Karen
Braga, Carolina
Charnalia, Abhishek
Ryan, Maeve
Rabbat, Mike
Drozdzal, Michal
Verbeek, Jakob
Romero-Soriano, Adriana
Computer Vision and Pattern Recognition
As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound, there are few unified benchmarking libraries that provide a framework for performing evaluations across many datasets and metrics. Furthermore, the rapid introduction of increasingly robust benchmarking methods requires that evaluation libraries remain flexible to new datasets and metrics. Finally, there remains a gap in synthesizing evaluations in order to deliver actionable takeaways about model performance. To enable unified, flexible, and actionable evaluations, we introduce EvalGIM (pronounced ''EvalGym''), a library for evaluating generative image models. EvalGIM contains broad support for datasets and metrics used to measure quality, diversity, and consistency of text-to-image generative models. In addition, EvalGIM is designed with flexibility for user customization as a top priority and contains a structure that allows plug-and-play additions of new datasets and metrics. To enable actionable evaluation insights, we introduce ''Evaluation Exercises'' that highlight takeaways for specific evaluation questions. The Evaluation Exercises contain easy-to-use and reproducible implementations of two state-of-the-art evaluation methods of text-to-image generative models: consistency-diversity-realism Pareto Fronts and disaggregated measurements of performance disparities across groups. EvalGIM also contains Evaluation Exercises that introduce two new analysis methods for text-to-image generative models: robustness analyses of model rankings and balanced evaluations across different prompt styles. We encourage text-to-image model exploration with EvalGIM and invite contributions at https://github.com/facebookresearch/EvalGIM/.
title EvalGIM: A Library for Evaluating Generative Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10604