Transformer Interpretability from Perspective of Attention and Gradient
Fuente:
arXiv
Salvato in:
| Autori principali: | Cui, Yongjin, Fan, Xiaohui, Chen, Huajun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Knowledge Model: Perspectives and Challenges
di: Chen, Huajun
Pubblicazione: (2023)
di: Chen, Huajun
Pubblicazione: (2023)
Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation
di: Cui, Yongjin, et al.
Pubblicazione: (2026)
di: Cui, Yongjin, et al.
Pubblicazione: (2026)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
di: Chen, Yihong, et al.
Pubblicazione: (2026)
di: Chen, Yihong, et al.
Pubblicazione: (2026)
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
di: Nam, Andrew, et al.
Pubblicazione: (2025)
di: Nam, Andrew, et al.
Pubblicazione: (2025)
SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
di: Bianco, Pedro Alejandro Dal, et al.
Pubblicazione: (2024)
di: Bianco, Pedro Alejandro Dal, et al.
Pubblicazione: (2024)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
di: El, Batu, et al.
Pubblicazione: (2025)
di: El, Batu, et al.
Pubblicazione: (2025)
Domain-Agnostic Molecular Generation with Chemical Feedback
di: Fang, Yin, et al.
Pubblicazione: (2023)
di: Fang, Yin, et al.
Pubblicazione: (2023)
GEGA: Graph Convolutional Networks and Evidence Retrieval Guided Attention for Enhanced Document-level Relation Extraction
di: Mao, Yanxu, et al.
Pubblicazione: (2024)
di: Mao, Yanxu, et al.
Pubblicazione: (2024)
Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
di: Dhayalkar, Sahil Rajesh
Pubblicazione: (2025)
di: Dhayalkar, Sahil Rajesh
Pubblicazione: (2025)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
di: Zixian, Wang
Pubblicazione: (2025)
di: Zixian, Wang
Pubblicazione: (2025)
Wireless Power Transfer and Intent-Driven Network Optimization in AAVs-assisted IoT for 6G Sustainable Connectivity
di: He, Xiaoming, et al.
Pubblicazione: (2025)
di: He, Xiaoming, et al.
Pubblicazione: (2025)
Lipschitz-aware Linearity Grafting for Certified Robustness
di: Han, Yongjin, et al.
Pubblicazione: (2025)
di: Han, Yongjin, et al.
Pubblicazione: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
di: Knights, Ethan
Pubblicazione: (2026)
di: Knights, Ethan
Pubblicazione: (2026)
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
di: Liao, Yi, et al.
Pubblicazione: (2025)
di: Liao, Yi, et al.
Pubblicazione: (2025)
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
di: Guan, Zhong, et al.
Pubblicazione: (2025)
di: Guan, Zhong, et al.
Pubblicazione: (2025)
On the Wasserstein Gradient Flow Interpretation of Drifting Models
di: Gretton, Arthur, et al.
Pubblicazione: (2026)
di: Gretton, Arthur, et al.
Pubblicazione: (2026)
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
di: Jafari, Aref, et al.
Pubblicazione: (2025)
di: Jafari, Aref, et al.
Pubblicazione: (2025)
CardioPatternFormer: Pattern-Guided Attention for Interpretable ECG Classification with Transformer Architecture
di: Uğraş, Berat Kutay, et al.
Pubblicazione: (2025)
di: Uğraş, Berat Kutay, et al.
Pubblicazione: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
di: Liang, Cheng, et al.
Pubblicazione: (2026)
di: Liang, Cheng, et al.
Pubblicazione: (2026)
Interpretable Vital Sign Forecasting with Model Agnostic Attention Maps
di: Liu, Yuwei, et al.
Pubblicazione: (2024)
di: Liu, Yuwei, et al.
Pubblicazione: (2024)
LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
di: Huang, Jihao, et al.
Pubblicazione: (2025)
di: Huang, Jihao, et al.
Pubblicazione: (2025)
Quantum Gradient Class Activation Map for Model Interpretability
di: Lin, Hsin-Yi, et al.
Pubblicazione: (2024)
di: Lin, Hsin-Yi, et al.
Pubblicazione: (2024)
Gradient-Based Program Synthesis with Neurally Interpreted Languages
di: Macfarlane, Matthew V., et al.
Pubblicazione: (2026)
di: Macfarlane, Matthew V., et al.
Pubblicazione: (2026)
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
di: Poupart, Yoann, et al.
Pubblicazione: (2025)
di: Poupart, Yoann, et al.
Pubblicazione: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
di: Yang, Yongjin, et al.
Pubblicazione: (2024)
di: Yang, Yongjin, et al.
Pubblicazione: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
di: Yoo, Haneul, et al.
Pubblicazione: (2024)
di: Yoo, Haneul, et al.
Pubblicazione: (2024)
STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models
di: Wang, Zerui, et al.
Pubblicazione: (2024)
di: Wang, Zerui, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
di: Bahador, Nooshin
Pubblicazione: (2025)
di: Bahador, Nooshin
Pubblicazione: (2025)
Emergence Transformer: Dynamical Temporal Attention Matters
di: Zhou, Zihan, et al.
Pubblicazione: (2026)
di: Zhou, Zihan, et al.
Pubblicazione: (2026)
Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model
di: Chen, Qifan, et al.
Pubblicazione: (2025)
di: Chen, Qifan, et al.
Pubblicazione: (2025)
Gradient Boosting within a Single Attention Layer
di: Sargolzaei, Saleh
Pubblicazione: (2026)
di: Sargolzaei, Saleh
Pubblicazione: (2026)
A Generalized Label Shift Perspective for Cross-Domain Gaze Estimation
di: Yang, Hao-Ran, et al.
Pubblicazione: (2025)
di: Yang, Hao-Ran, et al.
Pubblicazione: (2025)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
di: Ke, Yekun, et al.
Pubblicazione: (2024)
di: Ke, Yekun, et al.
Pubblicazione: (2024)
Closed-Form Interpretation of Neural Network Classifiers with Symbolic Gradients
di: Wetzel, Sebastian Johann
Pubblicazione: (2024)
di: Wetzel, Sebastian Johann
Pubblicazione: (2024)
Intrinsically Interpretable Attention via Sparse Post-Training
di: Draye, Florent, et al.
Pubblicazione: (2025)
di: Draye, Florent, et al.
Pubblicazione: (2025)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
di: Pratyush, Spandan
Pubblicazione: (2026)
di: Pratyush, Spandan
Pubblicazione: (2026)
Documenti analoghi
-
Large Knowledge Model: Perspectives and Challenges
di: Chen, Huajun
Pubblicazione: (2023) -
Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation
di: Cui, Yongjin, et al.
Pubblicazione: (2026) -
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
di: Chen, Yihong, et al.
Pubblicazione: (2026) -
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
di: Nam, Andrew, et al.
Pubblicazione: (2025) -
SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
di: Bianco, Pedro Alejandro Dal, et al.
Pubblicazione: (2024)