Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Donhauser, Konstantin, Arnal, Charles, Pezeshki, Mohammad, Cabannes, Vivien, Lopez-Paz, David, Ahuja, Kartik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Touring sampling with pushforward maps
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
Copyright-Protected Language Generation via Adaptive Model Fusion
by: Abad, Javier, et al.
Published: (2024)
by: Abad, Javier, et al.
Published: (2024)
Compositional Risk Minimization
by: Mahajan, Divyat, et al.
Published: (2024)
by: Mahajan, Divyat, et al.
Published: (2024)
On Provable Length and Compositional Generalization
by: Ahuja, Kartik, et al.
Published: (2024)
by: Ahuja, Kartik, et al.
Published: (2024)
Provable Benefits of In-Tool Learning for Large Language Models
by: Houliston, Sam, et al.
Published: (2025)
by: Houliston, Sam, et al.
Published: (2025)
Learning with Hidden Factorial Structure
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Mode Estimation with Partial Feedback
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026)
by: Arnal, Charles, et al.
Published: (2026)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
by: Goyal, Sachin, et al.
Published: (2025)
by: Goyal, Sachin, et al.
Published: (2025)
Clustering Head: A Visual Case Study of the Training Dynamics in Transformers
by: Odonnat, Ambroise, et al.
Published: (2024)
by: Odonnat, Ambroise, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
In-Context Learning through the Bayesian Prism
by: Panwar, Madhur, et al.
Published: (2023)
by: Panwar, Madhur, et al.
Published: (2023)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Brainformers: Trading Simplicity for Efficiency
by: Zhou, Yanqi, et al.
Published: (2023)
by: Zhou, Yanqi, et al.
Published: (2023)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
by: Gupta, Sharut, et al.
Published: (2026)
by: Gupta, Sharut, et al.
Published: (2026)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
by: Shyam, Vasudev, et al.
Published: (2024)
by: Shyam, Vasudev, et al.
Published: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
by: Sundaram, Shobhita, et al.
Published: (2026)
by: Sundaram, Shobhita, et al.
Published: (2026)
UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
by: Zhou, Lang, et al.
Published: (2026)
by: Zhou, Lang, et al.
Published: (2026)
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
by: Mahajan, Divyat, et al.
Published: (2025)
by: Mahajan, Divyat, et al.
Published: (2025)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
by: Chiang, Ting-Rui, et al.
Published: (2025)
by: Chiang, Ting-Rui, et al.
Published: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
by: Erden, Caner
Published: (2025)
by: Erden, Caner
Published: (2025)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
by: Lee, Changhun, et al.
Published: (2025)
by: Lee, Changhun, et al.
Published: (2025)
The Galerkin method beats Graph-Based Approaches for Spectral Algorithms
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
by: Sharma, Kartik, et al.
Published: (2026)
by: Sharma, Kartik, et al.
Published: (2026)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Similar Items
-
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024) -
Touring sampling with pushforward maps
by: Cabannes, Vivien, et al.
Published: (2023) -
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025) -
Copyright-Protected Language Generation via Adaptive Model Fusion
by: Abad, Javier, et al.
Published: (2024) -
Compositional Risk Minimization
by: Mahajan, Divyat, et al.
Published: (2024)