Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Nam, Andrew, Conklin, Henry, Yang, Yukang, Griffiths, Thomas, Cohen, Jonathan, Leslie, Sarah-Jane |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Task Representations in Neural Networks via Bayesian Ablation
by: Nam, Andrew, et al.
Published: (2025)
by: Nam, Andrew, et al.
Published: (2025)
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Representations as Language: An Information-Theoretic Framework for Interpretability
by: Conklin, Henry, et al.
Published: (2024)
by: Conklin, Henry, et al.
Published: (2024)
Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination
by: Jiang, Jinrui, et al.
Published: (2026)
by: Jiang, Jinrui, et al.
Published: (2026)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Information Structure in Mappings: An Approach to Learning, Representation, and Generalisation
by: Conklin, Henry
Published: (2025)
by: Conklin, Henry
Published: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
by: Liu, Qiawen Ella, et al.
Published: (2026)
by: Liu, Qiawen Ella, et al.
Published: (2026)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Heads or Tails: A Simple Example of Causal Abstractive Simulation
by: Simmons, Gabriel
Published: (2025)
by: Simmons, Gabriel
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
A Multi-Head Attention Soft Random Forest for Interpretable Patient No-Show Prediction
by: Amalina, Ninda Nurseha, et al.
Published: (2025)
by: Amalina, Ninda Nurseha, et al.
Published: (2025)
HeadRank: Decoding-Free Passage Reranking via Preference-Aligned Attention Heads
by: Wang, Juyuan, et al.
Published: (2026)
by: Wang, Juyuan, et al.
Published: (2026)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
by: Sandoval, Gustavo
Published: (2025)
by: Sandoval, Gustavo
Published: (2025)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
by: Ghodsi, Ali
Published: (2025)
by: Ghodsi, Ali
Published: (2025)
LATTE: Low-Precision Approximate Attention with Head-wise Trainable Threshold for Efficient Transformer
by: Wang, Jiing-Ping, et al.
Published: (2024)
by: Wang, Jiing-Ping, et al.
Published: (2024)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
Do Attention Heads Compete or Cooperate during Counting?
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
by: Mudarisov, Timur, et al.
Published: (2026)
by: Mudarisov, Timur, et al.
Published: (2026)
Analyzing Multi-Head Attention on Trojan BERT Models
by: Wang, Jingwei
Published: (2024)
by: Wang, Jingwei
Published: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers
by: Sun, Bohang, et al.
Published: (2025)
by: Sun, Bohang, et al.
Published: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
by: Sun, Xiangkun, et al.
Published: (2026)
by: Sun, Xiangkun, et al.
Published: (2026)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
by: Jin, Haotian, et al.
Published: (2025)
by: Jin, Haotian, et al.
Published: (2025)
Similar Items
-
Understanding Task Representations in Neural Networks via Bayesian Ablation
by: Nam, Andrew, et al.
Published: (2025) -
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025) -
Representations as Language: An Information-Theoretic Framework for Interpretability
by: Conklin, Henry, et al.
Published: (2024) -
Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination
by: Jiang, Jinrui, et al.
Published: (2026) -
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)