Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Magid, Salma Abdel, Pan, Weiwei, Warchol, Simon, Guo, Grace, Kim, Junsik, Rahman, Mahia, Pfister, Hanspeter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917219675406336
author Magid, Salma Abdel
Pan, Weiwei
Warchol, Simon
Guo, Grace
Kim, Junsik
Rahman, Mahia
Pfister, Hanspeter
author_facet Magid, Salma Abdel
Pan, Weiwei
Warchol, Simon
Guo, Grace
Kim, Junsik
Rahman, Mahia
Pfister, Hanspeter
contents Text-to-image (T2I) models are increasingly used in impactful real-life applications. As such, there is a growing need to audit these models to ensure that they generate desirable, task-appropriate images. However, systematically inspecting the associations between prompts and generated content in a human-understandable way remains challenging. To address this, we propose Concept2Concept, a framework where we characterize conditional distributions of vision language models using interpretable concepts and metrics that can be defined in terms of these concepts. This characterization allows us to use our framework to audit models and prompt-datasets. To demonstrate, we investigate several case studies of conditional distributions of prompts, such as user-defined distributions or empirical, real-world distributions. Lastly, we implement Concept2Concept as an open-source interactive visualization tool to facilitate use by non-technical end-users. A demo is available at https://tinyurl.com/Concept2ConceptDemo.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04634
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
Magid, Salma Abdel
Pan, Weiwei
Warchol, Simon
Guo, Grace
Kim, Junsik
Rahman, Mahia
Pfister, Hanspeter
Computer Vision and Pattern Recognition
Text-to-image (T2I) models are increasingly used in impactful real-life applications. As such, there is a growing need to audit these models to ensure that they generate desirable, task-appropriate images. However, systematically inspecting the associations between prompts and generated content in a human-understandable way remains challenging. To address this, we propose Concept2Concept, a framework where we characterize conditional distributions of vision language models using interpretable concepts and metrics that can be defined in terms of these concepts. This characterization allows us to use our framework to audit models and prompt-datasets. To demonstrate, we investigate several case studies of conditional distributions of prompts, such as user-defined distributions or empirical, real-world distributions. Lastly, we implement Concept2Concept as an open-source interactive visualization tool to facilitate use by non-technical end-users. A demo is available at https://tinyurl.com/Concept2ConceptDemo.
title Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.04634