GRADE: Quantifying Sample Diversity in Text-to-Image Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rassin, Royi, Slobodkin, Aviv, Ravfogel, Shauli, Elazar, Yanai, Goldberg, Yoav
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913729562542080
author Rassin, Royi
Slobodkin, Aviv
Ravfogel, Shauli
Elazar, Yanai
Goldberg, Yoav
author_facet Rassin, Royi
Slobodkin, Aviv
Ravfogel, Shauli
Elazar, Yanai
Goldberg, Yoav
contents We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and visual question-answering systems to identify relevant concept-specific axes of diversity (e.g., ``shape'' for the concept ``cookie''). It then estimates frequency distributions of concepts and their attributes and quantifies diversity using entropy. We use GRADE to measure the diversity of 12 models over a total of 720K images, revealing that all models display limited variation, with clear deterioration in stronger models. Further, we find that models often exhibit default behaviors, a phenomenon where a model consistently generates concepts with the same attributes (e.g., 98% of the cookies are round). Lastly, we show that a key reason for low diversity is underspecified captions in training data. Our work proposes an automatic, semantically-driven approach to measure sample diversity and highlights the stunning homogeneity in text-to-image models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22592
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GRADE: Quantifying Sample Diversity in Text-to-Image Models
Rassin, Royi
Slobodkin, Aviv
Ravfogel, Shauli
Elazar, Yanai
Goldberg, Yoav
Computer Vision and Pattern Recognition
We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and visual question-answering systems to identify relevant concept-specific axes of diversity (e.g., ``shape'' for the concept ``cookie''). It then estimates frequency distributions of concepts and their attributes and quantifies diversity using entropy. We use GRADE to measure the diversity of 12 models over a total of 720K images, revealing that all models display limited variation, with clear deterioration in stronger models. Further, we find that models often exhibit default behaviors, a phenomenon where a model consistently generates concepts with the same attributes (e.g., 98% of the cookies are round). Lastly, we show that a key reason for low diversity is underspecified captions in training data. Our work proposes an automatic, semantically-driven approach to measure sample diversity and highlights the stunning homogeneity in text-to-image models.
title GRADE: Quantifying Sample Diversity in Text-to-Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.22592