Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Toker, Michael, Orgad, Hadas, Ventura, Mor, Arad, Dana, Belinkov, Yonatan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913713699684352
author Toker, Michael
Orgad, Hadas
Ventura, Mor
Arad, Dana
Belinkov, Yonatan
author_facet Toker, Michael
Orgad, Hadas
Ventura, Mor
Arad, Dana
Belinkov, Yonatan
contents Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose the Diffusion Lens, a method for analyzing the text encoder of T2I models by generating images from its intermediate representations. Using the Diffusion Lens, we perform an extensive analysis of two recent T2I models. Exploring compound prompts, we find that complex scenes describing multiple objects are composed progressively and more slowly compared to simple scenes; Exploring knowledge retrieval, we find that representation of uncommon concepts requires further computation compared to common concepts, and that knowledge retrieval is gradual across layers. Overall, our findings provide valuable insights into the text encoder component in T2I pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05846
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
Toker, Michael
Orgad, Hadas
Ventura, Mor
Arad, Dana
Belinkov, Yonatan
Computer Vision and Pattern Recognition
Computation and Language
I.2.7; I.4.0
Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose the Diffusion Lens, a method for analyzing the text encoder of T2I models by generating images from its intermediate representations. Using the Diffusion Lens, we perform an extensive analysis of two recent T2I models. Exploring compound prompts, we find that complex scenes describing multiple objects are composed progressively and more slowly compared to simple scenes; Exploring knowledge retrieval, we find that representation of uncommon concepts requires further computation compared to common concepts, and that knowledge retrieval is gradual across layers. Overall, our findings provide valuable insights into the text encoder component in T2I pipelines.
title Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
topic Computer Vision and Pattern Recognition
Computation and Language
I.2.7; I.4.0
url https://arxiv.org/abs/2403.05846