Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms
Fuente:
arXiv
Saved in:
| Main Author: | Nair, Pravin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
by: Duhme, Christof, et al.
Published: (2026)
by: Duhme, Christof, et al.
Published: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Beyond $\ell_2$-norm and $\ell_\infty$-norm: A Curvature-Inspired $\ell_p$-Norm Scheme for Deep Neural Networks
by: Xu, Jianhao, et al.
Published: (2026)
by: Xu, Jianhao, et al.
Published: (2026)
Iterative Refinement for $\ell_p$-norm Regression
by: Adil, Deeksha, et al.
Published: (2019)
by: Adil, Deeksha, et al.
Published: (2019)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Optimal bounds for $\ell_p$ sensitivity sampling via $\ell_2$ augmentation
by: Munteanu, Alexander, et al.
Published: (2024)
by: Munteanu, Alexander, et al.
Published: (2024)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Online Correlation Clustering: Simultaneously Optimizing All $\ell_p$-norms
by: Davies, Sami, et al.
Published: (2025)
by: Davies, Sami, et al.
Published: (2025)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
On Traceability in $\ell_p$ Stochastic Convex Optimization
by: Voitovych, Sasha, et al.
Published: (2025)
by: Voitovych, Sasha, et al.
Published: (2025)
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
by: Overman, William, et al.
Published: (2026)
by: Overman, William, et al.
Published: (2026)
Coresets for Multiple $\ell_p$ Regression
by: Woodruff, David P., et al.
Published: (2024)
by: Woodruff, David P., et al.
Published: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Exploring the Impact of Temperature Scaling in Softmax for Classification and Adversarial Robustness
by: Xuan, Hao, et al.
Published: (2025)
by: Xuan, Hao, et al.
Published: (2025)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Optimality of Matrix Mechanism on $\ell_p^p$-metric
by: Liu, Jingcheng, et al.
Published: (2024)
by: Liu, Jingcheng, et al.
Published: (2024)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
by: Chou, Yuhong, et al.
Published: (2024)
by: Chou, Yuhong, et al.
Published: (2024)
DSL: Understanding and Improving Softmax Recommender Systems with Competition-Aware Scaling
by: Sahyouni, Bucher, et al.
Published: (2026)
by: Sahyouni, Bucher, et al.
Published: (2026)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
by: Indelman, Hedda Cohen, et al.
Published: (2024)
by: Indelman, Hedda Cohen, et al.
Published: (2024)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026)
by: Park, Kiho, et al.
Published: (2026)
Sharper Bounds for $\ell_p$ Sensitivity Sampling
by: Woodruff, David P., et al.
Published: (2023)
by: Woodruff, David P., et al.
Published: (2023)
Incentivized Lipschitz Bandits
by: Chakraborty, Sourav, et al.
Published: (2025)
by: Chakraborty, Sourav, et al.
Published: (2025)
Closed-form $\ell_r$ norm scaling with data for overparameterized linear regression and diagonal linear networks under $\ell_p$ bias
by: Zhang, Shuofeng, et al.
Published: (2025)
by: Zhang, Shuofeng, et al.
Published: (2025)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
by: Tao, Hongyi, et al.
Published: (2026)
by: Tao, Hongyi, et al.
Published: (2026)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023)
by: Kim, Hoyong, et al.
Published: (2023)
When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE
by: Erol, Melihcan, et al.
Published: (2026)
by: Erol, Melihcan, et al.
Published: (2026)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
1-Lipschitz Network Initialization for Certifiably Robust Classification Applications: A Decay Problem
by: Juston, Marius F. R., et al.
Published: (2025)
by: Juston, Marius F. R., et al.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Principles of Lipschitz continuity in neural networks
by: Luo, Róisín
Published: (2026)
by: Luo, Róisín
Published: (2026)
Similar Items
-
Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
by: Duhme, Christof, et al.
Published: (2026) -
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024) -
Beyond $\ell_2$-norm and $\ell_\infty$-norm: A Curvature-Inspired $\ell_p$-Norm Scheme for Deep Neural Networks
by: Xu, Jianhao, et al.
Published: (2026) -
Iterative Refinement for $\ell_p$-norm Regression
by: Adil, Deeksha, et al.
Published: (2019) -
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
by: Lu, Yao, et al.
Published: (2026)