Localizing Knowledge in Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zarei, Arman, Basu, Samyadeep, Rezaei, Keivan, Lin, Zihao, Nag, Sayan, Feizi, Soheil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
by: Balasubramanian, Sriram, et al.
Published: (2024)
by: Balasubramanian, Sriram, et al.
Published: (2024)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023)
by: Rezaei, Keivan, et al.
Published: (2023)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
Understanding Information Storage and Transfer in Multi-modal Large Language Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
DREW : Towards Robust Data Provenance by Leveraging Error-Controlled Watermarking
by: Saberi, Mehrdad, et al.
Published: (2024)
by: Saberi, Mehrdad, et al.
Published: (2024)
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
by: Saberi, Mehrdad, et al.
Published: (2023)
by: Saberi, Mehrdad, et al.
Published: (2023)
Understanding the Effect of using Semantically Meaningful Tokens for Visual Representation Learning
by: Kalibhat, Neha, et al.
Published: (2024)
by: Kalibhat, Neha, et al.
Published: (2024)
How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric Encodings
by: Audia, Samuel, et al.
Published: (2025)
by: Audia, Samuel, et al.
Published: (2025)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
by: Nag, Sayan, et al.
Published: (2024)
by: Nag, Sayan, et al.
Published: (2024)
Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs
by: Gandhi, Mona, et al.
Published: (2026)
by: Gandhi, Mona, et al.
Published: (2026)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Object-WIPER : Training-Free Object and Associated Effect Removal in Videos
by: Kushwaha, Saksham Singh, et al.
Published: (2026)
by: Kushwaha, Saksham Singh, et al.
Published: (2026)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
What do we learn from inverting CLIP models?
by: Kazemi, Hamid, et al.
Published: (2024)
by: Kazemi, Hamid, et al.
Published: (2024)
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
by: Hosseini, Parsa, et al.
Published: (2025)
by: Hosseini, Parsa, et al.
Published: (2025)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
VGDM: Vision-Guided Diffusion Model for Brain Tumor Detection and Segmentation
by: Behnam, Arman
Published: (2025)
by: Behnam, Arman
Published: (2025)
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
by: Saberi, Mehrdad, et al.
Published: (2026)
by: Saberi, Mehrdad, et al.
Published: (2026)
Weather-Aware Object Detection Transformer for Domain Adaptation
by: Gharatappeh, Soheil, et al.
Published: (2025)
by: Gharatappeh, Soheil, et al.
Published: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
by: Gan, Chaofan, et al.
Published: (2025)
by: Gan, Chaofan, et al.
Published: (2025)
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
Liveness Detection in Computer Vision: Transformer-based Self-Supervised Learning for Face Anti-Spoofing
by: Keresh, Arman, et al.
Published: (2024)
by: Keresh, Arman, et al.
Published: (2024)
Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
by: Huang, Sida, et al.
Published: (2025)
by: Huang, Sida, et al.
Published: (2025)
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
by: Zou, Xuechao, et al.
Published: (2025)
by: Zou, Xuechao, et al.
Published: (2025)
Training-Free Diffusion Framework for Stylized Image Generation with Identity Preservation
by: Rezaei, Mohammad Ali, et al.
Published: (2025)
by: Rezaei, Mohammad Ali, et al.
Published: (2025)
CardioDiT: Latent Diffusion Transformers for 4D Cardiac MRI Synthesis
by: Seyfarth, Marvin, et al.
Published: (2026)
by: Seyfarth, Marvin, et al.
Published: (2026)
MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models
by: Chowdhury, Sanjoy, et al.
Published: (2024)
by: Chowdhury, Sanjoy, et al.
Published: (2024)
Personalize Anything for Free with Diffusion Transformer
by: Feng, Haoran, et al.
Published: (2025)
by: Feng, Haoran, et al.
Published: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
by: Lin, Yijun, et al.
Published: (2025)
by: Lin, Yijun, et al.
Published: (2025)
Knowledge Distillation via the Target-aware Transformer
by: Lin, Sihao, et al.
Published: (2022)
by: Lin, Sihao, et al.
Published: (2022)
Masked Diffusion Captioning for Visual Feature Learning
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
JambaTalk: Speech-Driven 3D Talking Head Generation Based on Hybrid Transformer-Mamba Model
by: Jafari, Farzaneh, et al.
Published: (2024)
by: Jafari, Farzaneh, et al.
Published: (2024)
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
by: Song, Wenhui, et al.
Published: (2025)
by: Song, Wenhui, et al.
Published: (2025)
Similar Items
-
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
by: Zarei, Arman, et al.
Published: (2025) -
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024) -
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024) -
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
by: Balasubramanian, Sriram, et al.
Published: (2024) -
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)