Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jungwon, Ko, Jungmin, Byun, Dongnam, Rhee, Wonjong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
by: Park, Jungwon, et al.
Published: (2024)
by: Park, Jungwon, et al.
Published: (2024)
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
by: Byun, Dongnam, et al.
Published: (2025)
by: Byun, Dongnam, et al.
Published: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
by: Kim, Jimyeong, et al.
Published: (2024)
by: Kim, Jimyeong, et al.
Published: (2024)
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
by: Kim, Jimyeong, et al.
Published: (2025)
by: Kim, Jimyeong, et al.
Published: (2025)
FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models
by: So, Junhyuk, et al.
Published: (2023)
by: So, Junhyuk, et al.
Published: (2023)
Objective and Interpretable Breast Cosmesis Evaluation with Attention Guided Denoising Diffusion Anomaly Detection Model
by: Park, Sangjoon, et al.
Published: (2024)
by: Park, Sangjoon, et al.
Published: (2024)
Evaluating Feature Attribution Methods for Electrocardiogram
by: Suh, Jangwon, et al.
Published: (2022)
by: Suh, Jangwon, et al.
Published: (2022)
Enhancing Contrastive Learning with Efficient Combinatorial Positive Pairing
by: Kim, Jaeill, et al.
Published: (2024)
by: Kim, Jaeill, et al.
Published: (2024)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
by: Kim, Wonkyun, et al.
Published: (2024)
by: Kim, Wonkyun, et al.
Published: (2024)
I2AM: Interpreting Image-to-Image Latent Diffusion Models via Bi-Attribution Maps
by: Park, Junseo, et al.
Published: (2024)
by: Park, Junseo, et al.
Published: (2024)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
by: Byun, Ji Young, et al.
Published: (2025)
by: Byun, Ji Young, et al.
Published: (2025)
Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps
by: Lee, Seoyeon, et al.
Published: (2025)
by: Lee, Seoyeon, et al.
Published: (2025)
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
by: Liao, Yi, et al.
Published: (2025)
by: Liao, Yi, et al.
Published: (2025)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Fast Sampling Through The Reuse Of Attention Maps In Diffusion Models
by: Hunter, Rosco, et al.
Published: (2023)
by: Hunter, Rosco, et al.
Published: (2023)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging
by: Chung, Minjae, et al.
Published: (2025)
by: Chung, Minjae, et al.
Published: (2025)
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
by: Liu, Weimin, et al.
Published: (2026)
by: Liu, Weimin, et al.
Published: (2026)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
by: Kim, Mingyu, et al.
Published: (2025)
by: Kim, Mingyu, et al.
Published: (2025)
Simple Drop-in LoRA Conditioning on Attention Layers Will Improve Your Diffusion Model
by: Choi, Joo Young, et al.
Published: (2024)
by: Choi, Joo Young, et al.
Published: (2024)
LAMS-Edit: Latent and Attention Mixing with Schedulers for Improved Content Preservation in Diffusion-Based Image and Style Editing
by: Fu, Wingwa, et al.
Published: (2026)
by: Fu, Wingwa, et al.
Published: (2026)
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
by: Ko, Donggeun, et al.
Published: (2024)
by: Ko, Donggeun, et al.
Published: (2024)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
by: Jiang, Pengfei, et al.
Published: (2025)
by: Jiang, Pengfei, et al.
Published: (2025)
ReaMIL: Reasoning- and Evidence-Aware Multiple Instance Learning for Whole-Slide Histopathology
by: Jung, Hyun Do, et al.
Published: (2026)
by: Jung, Hyun Do, et al.
Published: (2026)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
by: Chen, Qiyuan, et al.
Published: (2026)
by: Chen, Qiyuan, et al.
Published: (2026)
Task-Agnostic Noisy Label Detection via Standardized Loss Aggregation
by: Park, Inhyuk, et al.
Published: (2026)
by: Park, Inhyuk, et al.
Published: (2026)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
by: Song, Yeji, et al.
Published: (2024)
by: Song, Yeji, et al.
Published: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video Reconstruction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
by: Park, Jungwon, et al.
Published: (2026)
by: Park, Jungwon, et al.
Published: (2026)
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
by: Zhao, Chenchen, et al.
Published: (2026)
by: Zhao, Chenchen, et al.
Published: (2026)
Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models
by: Lee, Tae-Young, et al.
Published: (2025)
by: Lee, Tae-Young, et al.
Published: (2025)
PQCAD-DM: Progressive Quantization and Calibration-Assisted Distillation for Extremely Efficient Diffusion Model
by: Ko, Beomseok, et al.
Published: (2025)
by: Ko, Beomseok, et al.
Published: (2025)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
by: Zou, Siyu, et al.
Published: (2024)
by: Zou, Siyu, et al.
Published: (2024)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
MAST: Mask-Guided Attention Mass Allocation for Training-Free Multi-Style Transfer
by: Kang, Dongkyung, et al.
Published: (2026)
by: Kang, Dongkyung, et al.
Published: (2026)
Similar Items
-
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
by: Park, Jungwon, et al.
Published: (2024) -
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
by: Byun, Dongnam, et al.
Published: (2025) -
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025) -
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026) -
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)