Saved in:
| Main Author: | Ran-Milo, Yuval |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.11487 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024)
by: Ran-Milo, Yuval, et al.
Published: (2024)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Do Neural Networks Need Gradient Descent to Generalize? A Theoretical Study
by: Alexander, Yotam, et al.
Published: (2025)
by: Alexander, Yotam, et al.
Published: (2025)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
by: İslamoğlu, Gamze, et al.
Published: (2023)
by: İslamoğlu, Gamze, et al.
Published: (2023)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
ASAP: Attention Sink Anchored Pruning
by: Lee, Jaehyuk, et al.
Published: (2026)
by: Lee, Jaehyuk, et al.
Published: (2026)
Mamba Knockout for Unraveling Factual Information Flow
by: Endy, Nir, et al.
Published: (2025)
by: Endy, Nir, et al.
Published: (2025)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
Provably Protecting Fine-Tuned LLMs from Training Data Extraction while Preserving Utility
by: Segal, Tom, et al.
Published: (2026)
by: Segal, Tom, et al.
Published: (2026)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
by: Mongaras, Gabriel, et al.
Published: (2025)
by: Mongaras, Gabriel, et al.
Published: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Stochastic Parroting in Temporal Attention -- Regulating the Diagonal Sink
by: Hankemeier, Victoria, et al.
Published: (2026)
by: Hankemeier, Victoria, et al.
Published: (2026)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
Provably Learning Attention with Queries
by: Bhattamishra, Satwik, et al.
Published: (2026)
by: Bhattamishra, Satwik, et al.
Published: (2026)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Softmax Transformers are Turing-Complete
by: Jiang, Hongjian, et al.
Published: (2025)
by: Jiang, Hongjian, et al.
Published: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
by: Hoffmann, David T., et al.
Published: (2023)
by: Hoffmann, David T., et al.
Published: (2023)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Similar Items
-
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
by: Ran-Milo, Yuval, et al.
Published: (2026) -
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026) -
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025) -
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024) -
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)