TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909734155583488 |
|---|---|
| author | Huang, Victor Shea-Jay Zhuo, Le Xin, Yi Wang, Zhaokai Wang, Fu-Yun Wang, Yuchi Zhang, Renrui Gao, Peng Li, Hongsheng |
| author_facet | Huang, Victor Shea-Jay Zhuo, Le Xin, Yi Wang, Zhaokai Wang, Fu-Yun Wang, Yuchi Zhang, Renrui Gao, Peng Li, Hongsheng |
| contents | Diffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE-Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs-a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_07050 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation Huang, Victor Shea-Jay Zhuo, Le Xin, Yi Wang, Zhaokai Wang, Fu-Yun Wang, Yuchi Zhang, Renrui Gao, Peng Li, Hongsheng Computer Vision and Pattern Recognition Artificial Intelligence Multimedia Diffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE-Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs-a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer. |
| title | TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Multimedia |
| url | https://arxiv.org/abs/2503.07050 |