Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
Fuente:
arXiv
Saved in:
| Main Author: | Sandoval, Gustavo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
by: Zhang, Jinhao, et al.
Published: (2026)
by: Zhang, Jinhao, et al.
Published: (2026)
Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery
by: Dudley, Carson, et al.
Published: (2025)
by: Dudley, Carson, et al.
Published: (2025)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
by: Hedström, Anna, et al.
Published: (2025)
by: Hedström, Anna, et al.
Published: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
An Odd Estimator for Shapley Values
by: Fumagalli, Fabian, et al.
Published: (2026)
by: Fumagalli, Fabian, et al.
Published: (2026)
Fly, Fail, Fix: Iterative Game Repair with Reinforcement Learning and Large Multimodal Models
by: Zook, Alex, et al.
Published: (2025)
by: Zook, Alex, et al.
Published: (2025)
Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis
by: Cai, Wang, et al.
Published: (2026)
by: Cai, Wang, et al.
Published: (2026)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
Do Attention Heads Compete or Cooperate during Counting?
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
by: Mudarisov, Timur, et al.
Published: (2026)
by: Mudarisov, Timur, et al.
Published: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
In-Context Algorithm Emulation in Fixed-Weight Transformers
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
by: Otsuka, Hikari, et al.
Published: (2025)
by: Otsuka, Hikari, et al.
Published: (2025)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
Causal Discovery in Linear Models with Unobserved Variables and Measurement Error
by: Yang, Yuqin, et al.
Published: (2024)
by: Yang, Yuqin, et al.
Published: (2024)
Contextual Counting: A Mechanistic Study of Transformers on a Quantitative Task
by: Golkar, Siavash, et al.
Published: (2024)
by: Golkar, Siavash, et al.
Published: (2024)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
by: Fatima, Sakina, et al.
Published: (2023)
by: Fatima, Sakina, et al.
Published: (2023)
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
by: Hsu, Hsin-Ling, et al.
Published: (2026)
by: Hsu, Hsin-Ling, et al.
Published: (2026)
Even-if Explanations: Formal Foundations, Priorities and Complexity
by: Alfano, Gianvincenzo, et al.
Published: (2024)
by: Alfano, Gianvincenzo, et al.
Published: (2024)
Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
by: Sato, Ryoma
Published: (2026)
by: Sato, Ryoma
Published: (2026)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
by: Wang, George, et al.
Published: (2024)
by: Wang, George, et al.
Published: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Quantum Error Mitigation with Attention Graph Transformers for Burgers Equation Solvers on NISQ Hardware
by: Tousi, Seyed Mohamad Ali, et al.
Published: (2025)
by: Tousi, Seyed Mohamad Ali, et al.
Published: (2025)
The Bayesian Geometry of Transformer Attention
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025)
by: Kimura, Daichi, et al.
Published: (2025)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
by: Freytes, Luis Rosario
Published: (2026)
by: Freytes, Luis Rosario
Published: (2026)
Similar Items
-
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
by: Bahador, Nooshin
Published: (2025) -
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
by: El, Batu, et al.
Published: (2025) -
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
by: Zhang, Jinhao, et al.
Published: (2026) -
Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery
by: Dudley, Carson, et al.
Published: (2025) -
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)