Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Rezazadeh, Navid, Davoodi, Arash Gholami |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
par: Rezazadeh, Navid, et autres
Publié: (2026)
par: Rezazadeh, Navid, et autres
Publié: (2026)
Softmax is not Enough (for Adaptive Conformal Classification)
par: Attar, Navid Akhavan, et autres
Publié: (2026)
par: Attar, Navid Akhavan, et autres
Publié: (2026)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
par: Gonsior, Julius, et autres
Publié: (2022)
par: Gonsior, Julius, et autres
Publié: (2022)
Softmax-free Linear Transformers
par: Lu, Jiachen, et autres
Publié: (2022)
par: Lu, Jiachen, et autres
Publié: (2022)
Universal Approximation with Softmax Attention
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
par: Lin, Zhixuan, et autres
Publié: (2025)
par: Lin, Zhixuan, et autres
Publié: (2025)
Scalable-Softmax Is Superior for Attention
par: Nakanishi, Ken M.
Publié: (2025)
par: Nakanishi, Ken M.
Publié: (2025)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
par: Cao, Yang, et autres
Publié: (2024)
par: Cao, Yang, et autres
Publié: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
par: Li, Yingru
Publié: (2025)
par: Li, Yingru
Publié: (2025)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
par: Collins, Liam, et autres
Publié: (2024)
par: Collins, Liam, et autres
Publié: (2024)
On the Invariants of Softmax Attention
par: Lee, Wonsuk
Publié: (2026)
par: Lee, Wonsuk
Publié: (2026)
The Information Geometry of Softmax: Probing and Steering
par: Park, Kiho, et autres
Publié: (2026)
par: Park, Kiho, et autres
Publié: (2026)
Softmax is not Enough (for Sharp Size Generalisation)
par: Veličković, Petar, et autres
Publié: (2024)
par: Veličković, Petar, et autres
Publié: (2024)
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
par: Overman, William, et autres
Publié: (2026)
par: Overman, William, et autres
Publié: (2026)
Fast Convergence of Softmax Policy Mirror Ascent
par: Asad, Reza, et autres
Publié: (2024)
par: Asad, Reza, et autres
Publié: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
par: Davoodi, Arash Gholami, et autres
Publié: (2024)
par: Davoodi, Arash Gholami, et autres
Publié: (2024)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
par: Davoodi, Arash Gholami, et autres
Publié: (2026)
par: Davoodi, Arash Gholami, et autres
Publié: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)
Exploring the Impact of Temperature Scaling in Softmax for Classification and Adversarial Robustness
par: Xuan, Hao, et autres
Publié: (2025)
par: Xuan, Hao, et autres
Publié: (2025)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
par: Hoffmann, David T., et autres
Publié: (2023)
par: Hoffmann, David T., et autres
Publié: (2023)
DSL: Understanding and Improving Softmax Recommender Systems with Competition-Aware Scaling
par: Sahyouni, Bucher, et autres
Publié: (2026)
par: Sahyouni, Bucher, et autres
Publié: (2026)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
par: Chou, Yuhong, et autres
Publié: (2024)
par: Chou, Yuhong, et autres
Publié: (2024)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
par: Indelman, Hedda Cohen, et autres
Publié: (2024)
par: Indelman, Hedda Cohen, et autres
Publié: (2024)
FLASH-D: FlashAttention with Hidden Softmax Division
par: Alexandridis, Kosmas, et autres
Publié: (2025)
par: Alexandridis, Kosmas, et autres
Publié: (2025)
When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE
par: Erol, Melihcan, et autres
Publié: (2026)
par: Erol, Melihcan, et autres
Publié: (2026)
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
par: Liu, Shiwei, et autres
Publié: (2024)
par: Liu, Shiwei, et autres
Publié: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
par: Sheen, Heejune, et autres
Publié: (2024)
par: Sheen, Heejune, et autres
Publié: (2024)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
par: Kim, Hoyong, et autres
Publié: (2023)
par: Kim, Hoyong, et autres
Publié: (2023)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
par: Zuo, Yifei, et autres
Publié: (2025)
par: Zuo, Yifei, et autres
Publié: (2025)
Softmax Linear Attention: Reclaiming Global Competition
par: Xu, Mingwei, et autres
Publié: (2026)
par: Xu, Mingwei, et autres
Publié: (2026)
PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation
par: Yang, Weiqin, et autres
Publié: (2024)
par: Yang, Weiqin, et autres
Publié: (2024)
Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms
par: Nair, Pravin
Publié: (2025)
par: Nair, Pravin
Publié: (2025)
Softmax gradient policy for variance minimization and risk-averse multi armed bandits
par: Turinici, Gabriel
Publié: (2026)
par: Turinici, Gabriel
Publié: (2026)
Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium
par: Lee-Jenkins, Christopher R.
Publié: (2025)
par: Lee-Jenkins, Christopher R.
Publié: (2025)
Making Sigmoid-MSE Great Again: Output Reset Challenges Softmax Cross-Entropy in Neural Network Classification
par: Tyagi, Kanishka, et autres
Publié: (2024)
par: Tyagi, Kanishka, et autres
Publié: (2024)
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
par: Kiselev, O. M.
Publié: (2026)
par: Kiselev, O. M.
Publié: (2026)
A Vulnerability of Attribution Methods Using Pre-Softmax Scores
par: Lerma, Miguel, et autres
Publié: (2023)
par: Lerma, Miguel, et autres
Publié: (2023)
Are Transformers More Robust? Towards Exact Robustness Verification for Transformers
par: Liao, Brian Hsuan-Cheng, et autres
Publié: (2022)
par: Liao, Brian Hsuan-Cheng, et autres
Publié: (2022)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
par: Yan, Fanqi, et autres
Publié: (2025)
par: Yan, Fanqi, et autres
Publié: (2025)
Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences
par: Koo, Jabin, et autres
Publié: (2026)
par: Koo, Jabin, et autres
Publié: (2026)
Documents similaires
-
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
par: Rezazadeh, Navid, et autres
Publié: (2026) -
Softmax is not Enough (for Adaptive Conformal Classification)
par: Attar, Navid Akhavan, et autres
Publié: (2026) -
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
par: Gonsior, Julius, et autres
Publié: (2022) -
Softmax-free Linear Transformers
par: Lu, Jiachen, et autres
Publié: (2022) -
Universal Approximation with Softmax Attention
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)