Localizing Knowledge in Diffusion Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zarei, Arman, Basu, Samyadeep, Rezaei, Keivan, Lin, Zihao, Nag, Sayan, Feizi, Soheil
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911398163906560
author Zarei, Arman
Basu, Samyadeep
Rezaei, Keivan
Lin, Zihao
Nag, Sayan
Feizi, Soheil
author_facet Zarei, Arman
Basu, Samyadeep
Rezaei, Keivan
Lin, Zihao
Nag, Sayan
Feizi, Soheil
contents Understanding how knowledge is distributed across the layers of generative models is crucial for improving interpretability, controllability, and adaptation. While prior work has explored knowledge localization in UNet-based architectures, Diffusion Transformer (DiT)-based models remain underexplored in this context. In this paper, we propose a model- and knowledge-agnostic method to localize where specific types of knowledge are encoded within the DiT blocks. We evaluate our method on state-of-the-art DiT-based models, including PixArt-alpha, FLUX, and SANA, across six diverse knowledge categories. We show that the identified blocks are both interpretable and causally linked to the expression of knowledge in generated outputs. Building on these insights, we apply our localization framework to two key applications: model personalization and knowledge unlearning. In both settings, our localized fine-tuning approach enables efficient and targeted updates, reducing computational cost, improving task-specific performance, and better preserving general model behavior with minimal interference to unrelated or surrounding content. Overall, our findings offer new insights into the internal structure of DiTs and introduce a practical pathway for more interpretable, efficient, and controllable model editing.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Localizing Knowledge in Diffusion Transformers
Zarei, Arman
Basu, Samyadeep
Rezaei, Keivan
Lin, Zihao
Nag, Sayan
Feizi, Soheil
Computer Vision and Pattern Recognition
Understanding how knowledge is distributed across the layers of generative models is crucial for improving interpretability, controllability, and adaptation. While prior work has explored knowledge localization in UNet-based architectures, Diffusion Transformer (DiT)-based models remain underexplored in this context. In this paper, we propose a model- and knowledge-agnostic method to localize where specific types of knowledge are encoded within the DiT blocks. We evaluate our method on state-of-the-art DiT-based models, including PixArt-alpha, FLUX, and SANA, across six diverse knowledge categories. We show that the identified blocks are both interpretable and causally linked to the expression of knowledge in generated outputs. Building on these insights, we apply our localization framework to two key applications: model personalization and knowledge unlearning. In both settings, our localized fine-tuning approach enables efficient and targeted updates, reducing computational cost, improving task-specific performance, and better preserving general model behavior with minimal interference to unrelated or surrounding content. Overall, our findings offer new insights into the internal structure of DiTs and introduce a practical pathway for more interpretable, efficient, and controllable model editing.
title Localizing Knowledge in Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18832