Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Seongmin, Hoover, Benjamin, Strobelt, Hendrik, Wang, Zijie J., Peng, ShengYun, Wright, Austin, Li, Kevin, Park, Haekyu, Yang, Haoyang, Chau, Duen Horng
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917764265934848
author Lee, Seongmin
Hoover, Benjamin
Strobelt, Hendrik
Wang, Zijie J.
Peng, ShengYun
Wright, Austin
Li, Kevin
Park, Haekyu
Yang, Haoyang
Chau, Duen Horng
author_facet Lee, Seongmin
Hoover, Benjamin
Strobelt, Hendrik
Wang, Zijie J.
Peng, ShengYun
Wright, Austin
Li, Kevin
Park, Haekyu
Yang, Haoyang
Chau, Duen Horng
contents Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion Explainer, the first interactive visualization tool that explains how Stable Diffusion transforms text prompts into images. Diffusion Explainer tightly integrates a visual overview of Stable Diffusion's complex structure with explanations of the underlying operations. By comparing image generation of prompt variants, users can discover the impact of keyword changes on image generation. A 56-participant user study demonstrates that Diffusion Explainer offers substantial learning benefits to non-experts. Our tool has been used by over 10,300 users from 124 countries at https://poloclub.github.io/diffusion-explainer/.
format Preprint
id arxiv_https___arxiv_org_abs_2305_03509
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
Lee, Seongmin
Hoover, Benjamin
Strobelt, Hendrik
Wang, Zijie J.
Peng, ShengYun
Wright, Austin
Li, Kevin
Park, Haekyu
Yang, Haoyang
Chau, Duen Horng
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion Explainer, the first interactive visualization tool that explains how Stable Diffusion transforms text prompts into images. Diffusion Explainer tightly integrates a visual overview of Stable Diffusion's complex structure with explanations of the underlying operations. By comparing image generation of prompt variants, users can discover the impact of keyword changes on image generation. A 56-participant user study demonstrates that Diffusion Explainer offers substantial learning benefits to non-experts. Our tool has been used by over 10,300 users from 124 countries at https://poloclub.github.io/diffusion-explainer/.
title Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2305.03509