Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeung, Calvin, Poduval, Prathyush, Zakeri, Ali, Zou, Zhuowen, Imani, Mohsen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910264132108288
author Yeung, Calvin
Poduval, Prathyush
Zakeri, Ali
Zou, Zhuowen
Imani, Mohsen
author_facet Yeung, Calvin
Poduval, Prathyush
Zakeri, Ali
Zou, Zhuowen
Imani, Mohsen
contents Text-to-image diffusion models generate images through an iterative denoising process, so internal neural layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable feature directions, but most approaches analyze activations at individual timesteps or condition on time rather than learning directly from full activation trajectories. In this work, we introduce residualized temporal SAEs for diffusion activation trajectories. We collect activations across denoising time, fit linear predictors between neighboring timesteps, and represent each trajectory using an initial activation together with residual components not explained by these linear dynamics. Training an SAE on this residualized representation encourages sparse latents to capture structure beyond what is linearly predictable. The residualized decoder directions can be mapped back into activation space, allowing each latent to be analyzed as a feature trajectory over denoising time. Through reconstruction and ablation studies, spatiotemporal feature analysis, and qualitative steering experiments on Stable Diffusion~1.5, we show that residualized temporal SAEs provide a useful framework for studying temporally structured diffusion activations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
Yeung, Calvin
Poduval, Prathyush
Zakeri, Ali
Zou, Zhuowen
Imani, Mohsen
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Text-to-image diffusion models generate images through an iterative denoising process, so internal neural layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable feature directions, but most approaches analyze activations at individual timesteps or condition on time rather than learning directly from full activation trajectories. In this work, we introduce residualized temporal SAEs for diffusion activation trajectories. We collect activations across denoising time, fit linear predictors between neighboring timesteps, and represent each trajectory using an initial activation together with residual components not explained by these linear dynamics. Training an SAE on this residualized representation encourages sparse latents to capture structure beyond what is linearly predictable. The residualized decoder directions can be mapped back into activation space, allowing each latent to be analyzed as a feature trajectory over denoising time. Through reconstruction and ablation studies, spatiotemporal feature analysis, and qualitative steering experiments on Stable Diffusion~1.5, we show that residualized temporal SAEs provide a useful framework for studying temporally structured diffusion activations.
title Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.27813