Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917764265934848 |
|---|---|
| author | Lee, Seongmin Hoover, Benjamin Strobelt, Hendrik Wang, Zijie J. Peng, ShengYun Wright, Austin Li, Kevin Park, Haekyu Yang, Haoyang Chau, Duen Horng |
| author_facet | Lee, Seongmin Hoover, Benjamin Strobelt, Hendrik Wang, Zijie J. Peng, ShengYun Wright, Austin Li, Kevin Park, Haekyu Yang, Haoyang Chau, Duen Horng |
| contents | Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion Explainer, the first interactive visualization tool that explains how Stable Diffusion transforms text prompts into images. Diffusion Explainer tightly integrates a visual overview of Stable Diffusion's complex structure with explanations of the underlying operations. By comparing image generation of prompt variants, users can discover the impact of keyword changes on image generation. A 56-participant user study demonstrates that Diffusion Explainer offers substantial learning benefits to non-experts. Our tool has been used by over 10,300 users from 124 countries at https://poloclub.github.io/diffusion-explainer/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_03509 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion Lee, Seongmin Hoover, Benjamin Strobelt, Hendrik Wang, Zijie J. Peng, ShengYun Wright, Austin Li, Kevin Park, Haekyu Yang, Haoyang Chau, Duen Horng Computation and Language Artificial Intelligence Human-Computer Interaction Machine Learning Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion Explainer, the first interactive visualization tool that explains how Stable Diffusion transforms text prompts into images. Diffusion Explainer tightly integrates a visual overview of Stable Diffusion's complex structure with explanations of the underlying operations. By comparing image generation of prompt variants, users can discover the impact of keyword changes on image generation. A 56-participant user study demonstrates that Diffusion Explainer offers substantial learning benefits to non-experts. Our tool has been used by over 10,300 users from 124 countries at https://poloclub.github.io/diffusion-explainer/. |
| title | Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion |
| topic | Computation and Language Artificial Intelligence Human-Computer Interaction Machine Learning |
| url | https://arxiv.org/abs/2305.03509 |