$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Dahye, Thomas, Xavier, Ghadiyaram, Deepti
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908476163227648
author Kim, Dahye
Thomas, Xavier
Ghadiyaram, Deepti
author_facet Kim, Dahye
Thomas, Xavier
Ghadiyaram, Deepti
contents We study $\textit{how}$ rich visual semantic information is represented within various layers and denoising timesteps of different diffusion architectures. We uncover monosemantic interpretable features by leveraging k-sparse autoencoders (k-SAE). We substantiate our mechanistic interpretations via transfer learning using light-weight classifiers on off-the-shelf diffusion models' features. On $4$ datasets, we demonstrate the effectiveness of diffusion features for representation learning. We provide an in-depth analysis of how different diffusion architectures, pre-training datasets, and language model conditioning impacts visual representation granularity, inductive biases, and transfer learning capabilities. Our work is a critical step towards deepening interpretability of black-box diffusion models. Code and visualizations available at: https://github.com/revelio-diffusion/revelio
format Preprint
id arxiv_https___arxiv_org_abs_2411_16725
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle $\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
Kim, Dahye
Thomas, Xavier
Ghadiyaram, Deepti
Computer Vision and Pattern Recognition
We study $\textit{how}$ rich visual semantic information is represented within various layers and denoising timesteps of different diffusion architectures. We uncover monosemantic interpretable features by leveraging k-sparse autoencoders (k-SAE). We substantiate our mechanistic interpretations via transfer learning using light-weight classifiers on off-the-shelf diffusion models' features. On $4$ datasets, we demonstrate the effectiveness of diffusion features for representation learning. We provide an in-depth analysis of how different diffusion architectures, pre-training datasets, and language model conditioning impacts visual representation granularity, inductive biases, and transfer learning capabilities. Our work is a critical step towards deepening interpretability of black-box diffusion models. Code and visualizations available at: https://github.com/revelio-diffusion/revelio
title $\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16725