GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Jo, Sehyeong, Jang, Gangjae, Park, Haesol |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decision-Aware Attention Propagation for Vision Transformer Explainability
by: Jo, Sehyeong, et al.
Published: (2026)
by: Jo, Sehyeong, et al.
Published: (2026)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
ComFe: An Interpretable Head for Vision Transformers
by: Mannix, Evelyn J., et al.
Published: (2024)
by: Mannix, Evelyn J., et al.
Published: (2024)
MAIR++: Improving Multi-view Attention Inverse Rendering with Implicit Lighting Representation
by: Choi, JunYong, et al.
Published: (2024)
by: Choi, JunYong, et al.
Published: (2024)
Multi-manifold Attention for Vision Transformers
by: Konstantinidis, Dimitrios, et al.
Published: (2022)
by: Konstantinidis, Dimitrios, et al.
Published: (2022)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
MHAFF: Multi-Head Attention Feature Fusion of CNN and Transformer for Cattle Identification
by: Dulal, Rabin, et al.
Published: (2025)
by: Dulal, Rabin, et al.
Published: (2025)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers
by: Bing, Zhaodong, et al.
Published: (2025)
by: Bing, Zhaodong, et al.
Published: (2025)
Interpretability-Aware Vision Transformer
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
Sparse but not Simpler: A Multi-Level Interpretability Analysis of Vision Transformers
by: Zhang, Siyu
Published: (2026)
by: Zhang, Siyu
Published: (2026)
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer
by: Sharma, Dhruv, et al.
Published: (2024)
by: Sharma, Dhruv, et al.
Published: (2024)
Representative Attention For Vision Transformers
by: Li, Yuntong, et al.
Published: (2026)
by: Li, Yuntong, et al.
Published: (2026)
Vision Transformers with Hierarchical Attention
by: Liu, Yun, et al.
Published: (2021)
by: Liu, Yun, et al.
Published: (2021)
IFQA: Interpretable Face Quality Assessment
by: Jo, Byungho, et al.
Published: (2022)
by: Jo, Byungho, et al.
Published: (2022)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
HAViT: Historical Attention Vision Transformer
by: Banik, Swarnendu, et al.
Published: (2026)
by: Banik, Swarnendu, et al.
Published: (2026)
Interactive Multi-Head Self-Attention with Linear Complexity
by: Kang, Hankyul, et al.
Published: (2024)
by: Kang, Hankyul, et al.
Published: (2024)
Multi-Head Attention Driven Dynamic Visual-Semantic Embedding for Enhanced Image-Text Matching
by: Chen, Wenjing
Published: (2024)
by: Chen, Wenjing
Published: (2024)
IG-FIQA: Improving Face Image Quality Assessment through Intra-class Variance Guidance robust to Inaccurate Pseudo-Labels
by: Kim, Minsoo, et al.
Published: (2024)
by: Kim, Minsoo, et al.
Published: (2024)
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
by: Liao, Yi, et al.
Published: (2025)
by: Liao, Yi, et al.
Published: (2025)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
Polyline Path Masked Attention for Vision Transformer
by: Zhao, Zhongchen, et al.
Published: (2025)
by: Zhao, Zhongchen, et al.
Published: (2025)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2025)
by: Garg, Mallika, et al.
Published: (2025)
V-NAW: Video-based Noise-aware Adaptive Weighting for Facial Expression Recognition
by: Lee, JunGyu, et al.
Published: (2025)
by: Lee, JunGyu, et al.
Published: (2025)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
by: Xiao, Chaodong, et al.
Published: (2026)
by: Xiao, Chaodong, et al.
Published: (2026)
SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
Interpretable Vision Transformers in Image Classification via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Multi Head Attention Enhanced Inception v3 for Cardiomegaly Detection
by: Karthik, Abishek, et al.
Published: (2025)
by: Karthik, Abishek, et al.
Published: (2025)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification
by: Gallée, Luisa, et al.
Published: (2025)
by: Gallée, Luisa, et al.
Published: (2025)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
Interpretable Vision Transformers in Monocular Depth Estimation via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
by: Ma, Chiyu, et al.
Published: (2024)
by: Ma, Chiyu, et al.
Published: (2024)
ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasets
by: Kim, Hoyoung, et al.
Published: (2026)
by: Kim, Hoyoung, et al.
Published: (2026)
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
by: Kwon, Minkyung, et al.
Published: (2025)
by: Kwon, Minkyung, et al.
Published: (2025)
Similar Items
-
Decision-Aware Attention Propagation for Vision Transformer Explainability
by: Jo, Sehyeong, et al.
Published: (2026) -
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024) -
ComFe: An Interpretable Head for Vision Transformers
by: Mannix, Evelyn J., et al.
Published: (2024) -
MAIR++: Improving Multi-view Attention Inverse Rendering with Implicit Lighting Representation
by: Choi, JunYong, et al.
Published: (2024) -
Multi-manifold Attention for Vision Transformers
by: Konstantinidis, Dimitrios, et al.
Published: (2022)