Do Attention Heads Compete or Cooperate during Counting?
Fuente:
arXiv
Saved in:
| Main Authors: | Zsámboki, Pál, Fraknói, Ádám, Gedeon, Máté, Kornai, András, Zombori, Zsolt |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Constraint-aware Learning of Probabilistic Sequential Models for Multi-Label Classification
by: Buleshnyi, Mykhailo, et al.
Published: (2025)
by: Buleshnyi, Mykhailo, et al.
Published: (2025)
Post-Norm can Resharpen Attention
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
Partial Label Learning for Automated Theorem Proving
by: Zombori, Zsolt, et al.
Published: (2025)
by: Zombori, Zsolt, et al.
Published: (2025)
A Comparative Analysis of Static Word Embeddings for Hungarian
by: Gedeon, Máté
Published: (2025)
by: Gedeon, Máté
Published: (2025)
Retrieval-Enhanced Few-Shot Prompting for Speech Event Extraction
by: Gedeon, Máté
Published: (2025)
by: Gedeon, Máté
Published: (2025)
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
by: Raj, Arjun, et al.
Published: (2024)
by: Raj, Arjun, et al.
Published: (2024)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
Vector Semantics
by: Kornai, András
Published: (2022)
by: Kornai, András
Published: (2022)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
Geometric Analysis of Token Selection in Multi-Head Attention
by: Mudarisov, Timur, et al.
Published: (2026)
by: Mudarisov, Timur, et al.
Published: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
by: Otsuka, Hikari, et al.
Published: (2025)
by: Otsuka, Hikari, et al.
Published: (2025)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
Scalable Context-Aware Graph Attention for Unsupervised Anomaly Detection in Large-Scale Mobile Networks
by: Malacarne, Sara, et al.
Published: (2026)
by: Malacarne, Sara, et al.
Published: (2026)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
by: Wang, George, et al.
Published: (2024)
by: Wang, George, et al.
Published: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Nonparametric Distribution Regression Re-calibration
by: Jung, Ádám, et al.
Published: (2026)
by: Jung, Ádám, et al.
Published: (2026)
Are Graph Attention Networks Able to Model Structural Information?
by: Noravesh, Farshad, et al.
Published: (2025)
by: Noravesh, Farshad, et al.
Published: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
by: Sandoval, Gustavo
Published: (2025)
by: Sandoval, Gustavo
Published: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Heterogeneous Multi-Agent Reinforcement Learning with Attention for Cooperative and Scalable Feature Transformation
by: Zhe, Tao, et al.
Published: (2025)
by: Zhe, Tao, et al.
Published: (2025)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
by: Ghodsi, Ali
Published: (2025)
by: Ghodsi, Ali
Published: (2025)
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
by: Urrutia, Felipe, et al.
Published: (2026)
by: Urrutia, Felipe, et al.
Published: (2026)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
Lemmas: Generation, Selection, Application
by: Rawson, Michael, et al.
Published: (2023)
by: Rawson, Michael, et al.
Published: (2023)
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
by: Jiang, Xinting, et al.
Published: (2026)
by: Jiang, Xinting, et al.
Published: (2026)
Verification of the Implicit World Model in a Generative Model via Adversarial Sequences
by: Balogh, András, et al.
Published: (2026)
by: Balogh, András, et al.
Published: (2026)
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Spatial Competence Benchmark
by: Vira, Jash, et al.
Published: (2026)
by: Vira, Jash, et al.
Published: (2026)
DiabetesNet: A Deep Learning Approach to Diabetes Diagnosis
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
A Multi-Modal CNN-LSTM Framework with Multi-Head Attention and Focal Loss for Real-Time Elderly Fall Detection
by: Zhou, Lijie, et al.
Published: (2026)
by: Zhou, Lijie, et al.
Published: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
Similar Items
-
Constraint-aware Learning of Probabilistic Sequential Models for Multi-Label Classification
by: Buleshnyi, Mykhailo, et al.
Published: (2025) -
Post-Norm can Resharpen Attention
by: Zsámboki, Pál, et al.
Published: (2025) -
Partial Label Learning for Automated Theorem Proving
by: Zombori, Zsolt, et al.
Published: (2025) -
A Comparative Analysis of Static Word Embeddings for Hungarian
by: Gedeon, Máté
Published: (2025) -
Retrieval-Enhanced Few-Shot Prompting for Speech Event Extraction
by: Gedeon, Máté
Published: (2025)