Self-Attention Decomposition For Training Free Diffusion Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anand, Tharun, Vali, Mohammad Hassan, Solin, Arno, Rosh, Green, Prasad, BH Pawan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911442921324544
author Anand, Tharun
Vali, Mohammad Hassan
Solin, Arno
Rosh, Green
Prasad, BH Pawan
author_facet Anand, Tharun
Vali, Mohammad Hassan
Solin, Arno
Rosh, Green
Prasad, BH Pawan
contents Diffusion models achieve remarkable fidelity in image synthesis, yet precise control over their outputs for targeted editing remains challenging. A key step toward controllability is to identify interpretable directions in the model's latent representations that correspond to semantic attributes. Existing approaches for finding interpretable directions typically rely on sampling large sets of images or training auxiliary networks, which limits efficiency. We propose an analytical method that derives semantic editing directions directly from the pretrained parameters of diffusion models, requiring neither additional data nor fine-tuning. Our insight is that self-attention weight matrices encode rich structural information about the data distribution learned during training. By computing the eigenvectors of these weight matrices, we obtain robust and interpretable editing directions. Experiments demonstrate that our method produces high-quality edits across multiple datasets while reducing editing time significantly by 60% over current benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Attention Decomposition For Training Free Diffusion Editing
Anand, Tharun
Vali, Mohammad Hassan
Solin, Arno
Rosh, Green
Prasad, BH Pawan
Computer Vision and Pattern Recognition
Diffusion models achieve remarkable fidelity in image synthesis, yet precise control over their outputs for targeted editing remains challenging. A key step toward controllability is to identify interpretable directions in the model's latent representations that correspond to semantic attributes. Existing approaches for finding interpretable directions typically rely on sampling large sets of images or training auxiliary networks, which limits efficiency. We propose an analytical method that derives semantic editing directions directly from the pretrained parameters of diffusion models, requiring neither additional data nor fine-tuning. Our insight is that self-attention weight matrices encode rich structural information about the data distribution learned during training. By computing the eigenvectors of these weight matrices, we obtain robust and interpretable editing directions. Experiments demonstrate that our method produces high-quality edits across multiple datasets while reducing editing time significantly by 60% over current benchmarks.
title Self-Attention Decomposition For Training Free Diffusion Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.22650