Geometric Analysis of Token Selection in Multi-Head Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mudarisov, Timur, Burtsev, Mikhal, Petrova, Tatiana, State, Radu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Limitations of Normalization in Attention Mechanism
von: Mudarisov, Timur, et al.
Veröffentlicht: (2025)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2025)
Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory
von: Petrova, Tatiana, et al.
Veröffentlicht: (2026)
von: Petrova, Tatiana, et al.
Veröffentlicht: (2026)
On the Volatility of Shapley-Based Contribution Metrics in Federated Learning
von: Geimer, Arno, et al.
Veröffentlicht: (2024)
von: Geimer, Arno, et al.
Veröffentlicht: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)
Position-Aware Sequential Attention for Accurate Next Item Recommendations
von: Nabiev, Timur, et al.
Veröffentlicht: (2026)
von: Nabiev, Timur, et al.
Veröffentlicht: (2026)
MoH: Multi-Head Attention as Mixture-of-Head Attention
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
von: Sagirova, Alsu, et al.
Veröffentlicht: (2025)
von: Sagirova, Alsu, et al.
Veröffentlicht: (2025)
Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
TransMLA: Multi-Head Latent Attention Is All You Need
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
Singular Vectors of Attention Heads Align with Features
von: Franco, Gabriel, et al.
Veröffentlicht: (2026)
von: Franco, Gabriel, et al.
Veröffentlicht: (2026)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
von: Julistiono, Addison Kristanto, et al.
Veröffentlicht: (2024)
von: Julistiono, Addison Kristanto, et al.
Veröffentlicht: (2024)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents
von: Petrova, Tatiana, et al.
Veröffentlicht: (2025)
von: Petrova, Tatiana, et al.
Veröffentlicht: (2025)
Do Attention Heads Compete or Cooperate during Counting?
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
von: Chepurova, Alla, et al.
Veröffentlicht: (2025)
von: Chepurova, Alla, et al.
Veröffentlicht: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
von: Freytes, Luis Rosario
Veröffentlicht: (2026)
von: Freytes, Luis Rosario
Veröffentlicht: (2026)
A Multi-Modal CNN-LSTM Framework with Multi-Head Attention and Focal Loss for Real-Time Elderly Fall Detection
von: Zhou, Lijie, et al.
Veröffentlicht: (2026)
von: Zhou, Lijie, et al.
Veröffentlicht: (2026)
MViewRouter: Internalizing Geometric Equivariance via Multi-view Alternating Attention for Combinatorial Routing
von: Liu, Shiyan, et al.
Veröffentlicht: (2026)
von: Liu, Shiyan, et al.
Veröffentlicht: (2026)
GraphT5: Unified Molecular Graph-Language Modeling via Multi-Modal Cross-Token Attention
von: Kim, Sangyeup, et al.
Veröffentlicht: (2025)
von: Kim, Sangyeup, et al.
Veröffentlicht: (2025)
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
von: Jiang, Xinting, et al.
Veröffentlicht: (2026)
von: Jiang, Xinting, et al.
Veröffentlicht: (2026)
T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning
von: Fu, Yanjun, et al.
Veröffentlicht: (2025)
von: Fu, Yanjun, et al.
Veröffentlicht: (2025)
ToMA: Token Merge with Attention for Diffusion Models
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
von: Bali, Karan, et al.
Veröffentlicht: (2026)
von: Bali, Karan, et al.
Veröffentlicht: (2026)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
von: Wang, George, et al.
Veröffentlicht: (2024)
von: Wang, George, et al.
Veröffentlicht: (2024)
Temporal-Spatial Tubelet Embedding for Cloud-Robust MSI Reconstruction using MSI-SAR Fusion: A Multi-Head Self-Attention Video Vision Transformer Approach
von: Wang, Yiqun, et al.
Veröffentlicht: (2025)
von: Wang, Yiqun, et al.
Veröffentlicht: (2025)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
von: Bu, Rui, et al.
Veröffentlicht: (2025)
von: Bu, Rui, et al.
Veröffentlicht: (2025)
A Multi-Head Attention Soft Random Forest for Interpretable Patient No-Show Prediction
von: Amalina, Ninda Nurseha, et al.
Veröffentlicht: (2025)
von: Amalina, Ninda Nurseha, et al.
Veröffentlicht: (2025)
Information as Structural Alignment: A Dynamical Theory of Continual Learning
von: Negulescu, Radu
Veröffentlicht: (2026)
von: Negulescu, Radu
Veröffentlicht: (2026)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
LongKey: Keyphrase Extraction for Long Documents
von: Alves, Jeovane Honorio, et al.
Veröffentlicht: (2024)
von: Alves, Jeovane Honorio, et al.
Veröffentlicht: (2024)
Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
von: Sandoval, Gustavo
Veröffentlicht: (2025)
von: Sandoval, Gustavo
Veröffentlicht: (2025)
A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases
von: Lombardo, Gianfranco, et al.
Veröffentlicht: (2026)
von: Lombardo, Gianfranco, et al.
Veröffentlicht: (2026)
Parkinson's Disease Detection from Resting State EEG using Multi-Head Graph Structure Learning with Gradient Weighted Graph Attention Explanations
von: Neves, Christopher, et al.
Veröffentlicht: (2024)
von: Neves, Christopher, et al.
Veröffentlicht: (2024)
Multi-Head Attention Is a Multi-Player Game
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
Enhancing Customer Service Chatbots with Context-Aware NLU through Selective Attention and Multi-task Learning
von: Nandi, Subhadip, et al.
Veröffentlicht: (2025)
von: Nandi, Subhadip, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Limitations of Normalization in Attention Mechanism
von: Mudarisov, Timur, et al.
Veröffentlicht: (2025) -
Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory
von: Petrova, Tatiana, et al.
Veröffentlicht: (2026) -
On the Volatility of Shapley-Based Contribution Metrics in Federated Learning
von: Geimer, Arno, et al.
Veröffentlicht: (2024) -
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024) -
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)