DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Simone, Zoe, Boggust, Angie, Satyanarayan, Arvind, Wilson, Ashia
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913222706069504
author De Simone, Zoe
Boggust, Angie
Satyanarayan, Arvind
Wilson, Ashia
author_facet De Simone, Zoe
Boggust, Angie
Satyanarayan, Arvind
Wilson, Ashia
contents Generative text-to-image (TTI) models produce high-quality images from short textual descriptions and are widely used in academic and creative domains. Like humans, TTI models have a worldview, a conception of the world learned from their training data and task that influences the images they generate for a given prompt. However, the worldviews of TTI models are often hidden from users, making it challenging for users to build intuition about TTI outputs, and they are often misaligned with users' worldviews, resulting in output images that do not match user expectations. In response, we introduce DiffusionWorldViewer, an interactive interface that exposes a TTI model's worldview across output demographics and provides editing tools for aligning output images with user perspectives. In a user study with 18 diverse TTI users, we find that DiffusionWorldViewer helps users represent their varied viewpoints in generated images and challenge the limited worldview reflected in current TTI models.
format Preprint
id arxiv_https___arxiv_org_abs_2309_09944
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
De Simone, Zoe
Boggust, Angie
Satyanarayan, Arvind
Wilson, Ashia
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
Generative text-to-image (TTI) models produce high-quality images from short textual descriptions and are widely used in academic and creative domains. Like humans, TTI models have a worldview, a conception of the world learned from their training data and task that influences the images they generate for a given prompt. However, the worldviews of TTI models are often hidden from users, making it challenging for users to build intuition about TTI outputs, and they are often misaligned with users' worldviews, resulting in output images that do not match user expectations. In response, we introduce DiffusionWorldViewer, an interactive interface that exposes a TTI model's worldview across output demographics and provides editing tools for aligning output images with user perspectives. In a user study with 18 diverse TTI users, we find that DiffusionWorldViewer helps users represent their varied viewpoints in generated images and challenge the limited worldview reflected in current TTI models.
title DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2309.09944