Interpreting Transformers Through Attention Head Intervention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kadem, Mason, Zheng, Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Human-Centered Ambient and Wearable Sensing for Automated Monitoring in Dementia Care: A Scoping Review
von: Kadem, Mason, et al.
Veröffentlicht: (2026)
von: Kadem, Mason, et al.
Veröffentlicht: (2026)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Surgical Repair of Collapsed Attention Heads in ALiBi Transformers
von: Schallon, Palmer
Veröffentlicht: (2026)
von: Schallon, Palmer
Veröffentlicht: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
Knocking-Heads Attention
von: Zhou, Zhanchao, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanchao, et al.
Veröffentlicht: (2025)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
von: Musat, Tiberiu
Veröffentlicht: (2024)
von: Musat, Tiberiu
Veröffentlicht: (2024)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
von: Basile, Lorenzo, et al.
Veröffentlicht: (2025)
von: Basile, Lorenzo, et al.
Veröffentlicht: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
Attention Heads of Large Language Models: A Survey
von: Zheng, Zifan, et al.
Veröffentlicht: (2024)
von: Zheng, Zifan, et al.
Veröffentlicht: (2024)
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads
von: Yang, Yi, et al.
Veröffentlicht: (2023)
von: Yang, Yi, et al.
Veröffentlicht: (2023)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
Perception Is All You Need: A Neuroscience Framework for Low Cost Sensorless Gaze in HRI
von: Kadem, Mason
Veröffentlicht: (2026)
von: Kadem, Mason
Veröffentlicht: (2026)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
von: Bianco, Pedro Alejandro Dal, et al.
Veröffentlicht: (2024)
von: Bianco, Pedro Alejandro Dal, et al.
Veröffentlicht: (2024)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
von: Lin, Xihui, et al.
Veröffentlicht: (2024)
von: Lin, Xihui, et al.
Veröffentlicht: (2024)
Inferring Functionality of Attention Heads from their Parameters
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
Head-wise Shareable Attention for Large Language Models
von: Cao, Zouying, et al.
Veröffentlicht: (2024)
von: Cao, Zouying, et al.
Veröffentlicht: (2024)
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
von: Hua, Kai, et al.
Veröffentlicht: (2025)
von: Hua, Kai, et al.
Veröffentlicht: (2025)
Sycophancy Hides Linearly in the Attention Heads
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
Interpreting Negation in GPT-2: Layer- and Head-Level Causal Analysis
von: Mofael, Abdullah Al, et al.
Veröffentlicht: (2026)
von: Mofael, Abdullah Al, et al.
Veröffentlicht: (2026)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
von: Ho, Zheng Yi, et al.
Veröffentlicht: (2024)
von: Ho, Zheng Yi, et al.
Veröffentlicht: (2024)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
von: Ji, Tao, et al.
Veröffentlicht: (2025)
von: Ji, Tao, et al.
Veröffentlicht: (2025)
When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-Incrementality
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2025)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2025)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
von: Ma, Youmi, et al.
Veröffentlicht: (2026)
von: Ma, Youmi, et al.
Veröffentlicht: (2026)
Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
von: Zhang, Weixu, et al.
Veröffentlicht: (2026)
von: Zhang, Weixu, et al.
Veröffentlicht: (2026)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
von: Kiruluta, Andrew
Veröffentlicht: (2025)
von: Kiruluta, Andrew
Veröffentlicht: (2025)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
von: Han, Tingxu, et al.
Veröffentlicht: (2025)
von: Han, Tingxu, et al.
Veröffentlicht: (2025)
Analyzing Multi-Head Attention on Trojan BERT Models
von: Wang, Jingwei
Veröffentlicht: (2024)
von: Wang, Jingwei
Veröffentlicht: (2024)
CHAI: Clustered Head Attention for Efficient LLM Inference
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
Latent Multi-Head Attention for Small Language Models
von: Mehta, Sushant, et al.
Veröffentlicht: (2025)
von: Mehta, Sushant, et al.
Veröffentlicht: (2025)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
von: Wang, Yian, et al.
Veröffentlicht: (2026)
von: Wang, Yian, et al.
Veröffentlicht: (2026)
A Transformer with Stack Attention
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel
von: Yamagiwa, Hiroaki, et al.
Veröffentlicht: (2026)
von: Yamagiwa, Hiroaki, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Human-Centered Ambient and Wearable Sensing for Automated Monitoring in Dementia Care: A Scoping Review
von: Kadem, Mason, et al.
Veröffentlicht: (2026) -
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025) -
Surgical Repair of Collapsed Attention Heads in ALiBi Transformers
von: Schallon, Palmer
Veröffentlicht: (2026) -
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024) -
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)