Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Pu, Yuanhao, Lian, Defu, Chen, Enhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025)
by: Pu, Yuanhao, et al.
Published: (2025)
Beyond Surrogates: A Quantitative Analysis for Inter-Metric Relationships
by: Pu, Yuanhao, et al.
Published: (2026)
by: Pu, Yuanhao, et al.
Published: (2026)
Adaptive Sampled Softmax with Inverted Multi-Index: Methods, Theory and Applications
by: Chen, Jin, et al.
Published: (2025)
by: Chen, Jin, et al.
Published: (2025)
PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation
by: Yang, Weiqin, et al.
Published: (2024)
by: Yang, Weiqin, et al.
Published: (2024)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier
by: Hayta, Berk, et al.
Published: (2026)
by: Hayta, Berk, et al.
Published: (2026)
CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption
by: Park, Hanjun, et al.
Published: (2026)
by: Park, Hanjun, et al.
Published: (2026)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Softmax as a Lagrangian-Legendrian Seam
by: Lee-Jenkins, Christopher R.
Published: (2025)
by: Lee-Jenkins, Christopher R.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
by: Hoffmann, Marcel, et al.
Published: (2025)
by: Hoffmann, Marcel, et al.
Published: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
Softmax pin to Cauchy-Poisson Study
by: Jones, Gareth
Published: (2026)
by: Jones, Gareth
Published: (2026)
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026)
by: Park, Kiho, et al.
Published: (2026)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
Softmax Transformers are Turing-Complete
by: Jiang, Hongjian, et al.
Published: (2025)
by: Jiang, Hongjian, et al.
Published: (2025)
Online Continual Learning via Logit Adjusted Softmax
by: Huang, Zhehao, et al.
Published: (2023)
by: Huang, Zhehao, et al.
Published: (2023)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
by: Mongaras, Gabriel, et al.
Published: (2025)
by: Mongaras, Gabriel, et al.
Published: (2025)
Beyond Softmax: A Natural Parameterization for Categorical Random Variables
by: Manenti, Alessandro, et al.
Published: (2025)
by: Manenti, Alessandro, et al.
Published: (2025)
Beyond Softmax: A New Perspective on Gradient Bandits
by: Melo, Emerson, et al.
Published: (2025)
by: Melo, Emerson, et al.
Published: (2025)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
by: Klein, Sara, et al.
Published: (2023)
by: Klein, Sara, et al.
Published: (2023)
Efficient Machine Unlearning via Influence Approximation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Binary Hypothesis Testing for Softmax Models and Leverage Score Models
by: Gu, Yuzhou, et al.
Published: (2024)
by: Gu, Yuzhou, et al.
Published: (2024)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Similar Items
-
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025) -
Beyond Surrogates: A Quantitative Analysis for Inter-Metric Relationships
by: Pu, Yuanhao, et al.
Published: (2026) -
Adaptive Sampled Softmax with Inverted Multi-Index: Methods, Theory and Applications
by: Chen, Jin, et al.
Published: (2025) -
PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation
by: Yang, Weiqin, et al.
Published: (2024) -
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)