Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, George, Hoogland, Jesse, van Wingerden, Stan, Furman, Zach, Murfet, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
by: Lau, Edmund, et al.
Published: (2023)
by: Lau, Edmund, et al.
Published: (2023)
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024)
by: Hoogland, Jesse, et al.
Published: (2024)
The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
by: Adam, Maxwell, et al.
Published: (2025)
by: Adam, Maxwell, et al.
Published: (2025)
Structural Inference: Interpreting Small Language Models with Susceptibilities
by: Baker, Garrett, et al.
Published: (2025)
by: Baker, Garrett, et al.
Published: (2025)
Towards Spectroscopy: Susceptibility Clusters in Language Models
by: Gordon, Andrew, et al.
Published: (2026)
by: Gordon, Andrew, et al.
Published: (2026)
Estimating the Local Learning Coefficient at Scale
by: Furman, Zach, et al.
Published: (2024)
by: Furman, Zach, et al.
Published: (2024)
Dynamics of Transient Structure in In-Context Linear Regression Transformers
by: Carroll, Liam, et al.
Published: (2025)
by: Carroll, Liam, et al.
Published: (2025)
Modes of Sequence Models and Learning Coefficients
by: Chen, Zhongtian, et al.
Published: (2025)
by: Chen, Zhongtian, et al.
Published: (2025)
Bayesian Influence Functions for Hessian-Free Data Attribution
by: Kreer, Philipp Alexander, et al.
Published: (2025)
by: Kreer, Philipp Alexander, et al.
Published: (2025)
Graph-level Protein Representation Learning by Structure Knowledge Refinement
by: Wang, Ge, et al.
Published: (2024)
by: Wang, Ge, et al.
Published: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
From Global to Local: A Scalable Benchmark for Local Posterior Sampling
by: Hitchcock, Rohan, et al.
Published: (2025)
by: Hitchcock, Rohan, et al.
Published: (2025)
Are Foundation Models Useful for Bankruptcy Prediction?
by: Kostrzewa, Marcin, et al.
Published: (2025)
by: Kostrzewa, Marcin, et al.
Published: (2025)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Deep Fusion: Capturing Dependencies in Contrastive Learning via Transformer Projection Heads
by: Li, Huanran, et al.
Published: (2024)
by: Li, Huanran, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
Grouped Differential Attention
by: Lim, Junghwan, et al.
Published: (2025)
by: Lim, Junghwan, et al.
Published: (2025)
Offline Reinforcement Learning for Learning to Dispatch for Job Shop Scheduling
by: van Remmerden, Jesse, et al.
Published: (2024)
by: van Remmerden, Jesse, et al.
Published: (2024)
Scalable Meta-Learning via Mixed-Mode Differentiation
by: Kemaev, Iurii, et al.
Published: (2025)
by: Kemaev, Iurii, et al.
Published: (2025)
Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Data
by: van Remmerden, Jesse, et al.
Published: (2025)
by: van Remmerden, Jesse, et al.
Published: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Geometric Analysis of Token Selection in Multi-Head Attention
by: Mudarisov, Timur, et al.
Published: (2026)
by: Mudarisov, Timur, et al.
Published: (2026)
Do Attention Heads Compete or Cooperate during Counting?
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
Unifying Perspectives: Plausible Counterfactual Explanations on Global, Group-wise, and Local Levels
by: Furman, Oleksii, et al.
Published: (2024)
by: Furman, Oleksii, et al.
Published: (2024)
AttentionSmithy: A Modular Framework for Rapid Transformer Development and Customization
by: Cranney, Caleb, et al.
Published: (2025)
by: Cranney, Caleb, et al.
Published: (2025)
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
by: Urrutia, Felipe, et al.
Published: (2026)
by: Urrutia, Felipe, et al.
Published: (2026)
Patterning: The Dual of Interpretability
by: Wang, George, et al.
Published: (2026)
by: Wang, George, et al.
Published: (2026)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
by: Otsuka, Hikari, et al.
Published: (2025)
by: Otsuka, Hikari, et al.
Published: (2025)
Deep Learning for Modeling and Dispatching Hybrid Wind Farm Power Generation
by: Lawrence, Zach, et al.
Published: (2025)
by: Lawrence, Zach, et al.
Published: (2025)
Learning to Model Graph Structural Information on MLPs via Graph Structure Self-Contrasting
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Design Principles for Sequence Models via Coefficient Dynamics
by: Sieber, Jerome, et al.
Published: (2025)
by: Sieber, Jerome, et al.
Published: (2025)
CCPL: Cross-modal Contrastive Protein Learning
by: Zheng, Jiangbin, et al.
Published: (2023)
by: Zheng, Jiangbin, et al.
Published: (2023)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Applying Deep Learning to Anomaly Detection of Russian Satellite Activity for Indications Prior to Military Activity
by: Kurtenbach, David, et al.
Published: (2025)
by: Kurtenbach, David, et al.
Published: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Generalized Learning of Coefficients in Spectral Graph Convolutional Networks
by: Coşkun, Mustafa, et al.
Published: (2024)
by: Coşkun, Mustafa, et al.
Published: (2024)
Similar Items
-
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
by: Lau, Edmund, et al.
Published: (2023) -
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
by: Urdshals, Einar, et al.
Published: (2025) -
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024) -
The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
by: Adam, Maxwell, et al.
Published: (2025) -
Structural Inference: Interpreting Small Language Models with Susceptibilities
by: Baker, Garrett, et al.
Published: (2025)