GRADE: Quantifying Sample Diversity in Text-to-Image Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913729562542080 |
|---|---|
| author | Rassin, Royi Slobodkin, Aviv Ravfogel, Shauli Elazar, Yanai Goldberg, Yoav |
| author_facet | Rassin, Royi Slobodkin, Aviv Ravfogel, Shauli Elazar, Yanai Goldberg, Yoav |
| contents | We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and visual question-answering systems to identify relevant concept-specific axes of diversity (e.g., ``shape'' for the concept ``cookie''). It then estimates frequency distributions of concepts and their attributes and quantifies diversity using entropy. We use GRADE to measure the diversity of 12 models over a total of 720K images, revealing that all models display limited variation, with clear deterioration in stronger models. Further, we find that models often exhibit default behaviors, a phenomenon where a model consistently generates concepts with the same attributes (e.g., 98% of the cookies are round). Lastly, we show that a key reason for low diversity is underspecified captions in training data. Our work proposes an automatic, semantically-driven approach to measure sample diversity and highlights the stunning homogeneity in text-to-image models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_22592 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GRADE: Quantifying Sample Diversity in Text-to-Image Models Rassin, Royi Slobodkin, Aviv Ravfogel, Shauli Elazar, Yanai Goldberg, Yoav Computer Vision and Pattern Recognition We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and visual question-answering systems to identify relevant concept-specific axes of diversity (e.g., ``shape'' for the concept ``cookie''). It then estimates frequency distributions of concepts and their attributes and quantifies diversity using entropy. We use GRADE to measure the diversity of 12 models over a total of 720K images, revealing that all models display limited variation, with clear deterioration in stronger models. Further, we find that models often exhibit default behaviors, a phenomenon where a model consistently generates concepts with the same attributes (e.g., 98% of the cookies are round). Lastly, we show that a key reason for low diversity is underspecified captions in training data. Our work proposes an automatic, semantically-driven approach to measure sample diversity and highlights the stunning homogeneity in text-to-image models. |
| title | GRADE: Quantifying Sample Diversity in Text-to-Image Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.22592 |