Softmax as a Lagrangian-Legendrian Seam
Fuente:
arXiv
Saved in:
| Main Author: | Lee-Jenkins, Christopher R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium
by: Lee-Jenkins, Christopher R.
Published: (2025)
by: Lee-Jenkins, Christopher R.
Published: (2025)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption
by: Park, Hanjun, et al.
Published: (2026)
by: Park, Hanjun, et al.
Published: (2026)
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
by: Pu, Yuanhao, et al.
Published: (2026)
by: Pu, Yuanhao, et al.
Published: (2026)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
by: Hoffmann, Marcel, et al.
Published: (2025)
by: Hoffmann, Marcel, et al.
Published: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025)
by: Pu, Yuanhao, et al.
Published: (2025)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
Online Continual Learning via Logit Adjusted Softmax
by: Huang, Zhehao, et al.
Published: (2023)
by: Huang, Zhehao, et al.
Published: (2023)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
by: Mongaras, Gabriel, et al.
Published: (2025)
by: Mongaras, Gabriel, et al.
Published: (2025)
Beyond Softmax: A Natural Parameterization for Categorical Random Variables
by: Manenti, Alessandro, et al.
Published: (2025)
by: Manenti, Alessandro, et al.
Published: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
Binary Hypothesis Testing for Softmax Models and Leverage Score Models
by: Gu, Yuzhou, et al.
Published: (2024)
by: Gu, Yuzhou, et al.
Published: (2024)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Softlog-Softmax Layers and Divergences Contribute to a Computationally Dependable Ensemble Learning
by: Atto, Abdourrahmane Mahamane
Published: (2025)
by: Atto, Abdourrahmane Mahamane
Published: (2025)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Probability Bounding: Post-Hoc Calibration via Box-Constrained Softmax
by: Atarashi, Kyohei, et al.
Published: (2025)
by: Atarashi, Kyohei, et al.
Published: (2025)
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
by: Masarczyk, Wojciech, et al.
Published: (2025)
by: Masarczyk, Wojciech, et al.
Published: (2025)
Adaptive Sampled Softmax with Inverted Multi-Index: Methods, Theory and Applications
by: Chen, Jin, et al.
Published: (2025)
by: Chen, Jin, et al.
Published: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Improving Explainability of Softmax Classifiers Using a Prototype-Based Joint Embedding Method
by: Sit, Hilarie, et al.
Published: (2024)
by: Sit, Hilarie, et al.
Published: (2024)
Softmax pin to Cauchy-Poisson Study
by: Jones, Gareth
Published: (2026)
by: Jones, Gareth
Published: (2026)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
Similar Items
-
Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium
by: Lee-Jenkins, Christopher R.
Published: (2025) -
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025) -
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026) -
CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption
by: Park, Hanjun, et al.
Published: (2026) -
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
by: Pu, Yuanhao, et al.
Published: (2026)