Old Optimizer, New Norm: An Anthology
Fuente:
arXiv
Saved in:
| Main Authors: | Bernstein, Jeremy, Newhouse, Laker |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modular Duality in Deep Learning
by: Bernstein, Jeremy, et al.
Published: (2024)
by: Bernstein, Jeremy, et al.
Published: (2024)
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
by: Liu, Zijian
Published: (2025)
by: Liu, Zijian
Published: (2025)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Small Gradient Norm Regret for Online Convex Optimization
by: Gao, Wenzhi, et al.
Published: (2026)
by: Gao, Wenzhi, et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise
by: Zhang, Jiayu, et al.
Published: (2026)
by: Zhang, Jiayu, et al.
Published: (2026)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
by: Kravatskiy, Alexey, et al.
Published: (2025)
by: Kravatskiy, Alexey, et al.
Published: (2025)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Universal Architectures for the Learning of Polyhedral Norms and Convex Regularizers
by: Unser, Michael, et al.
Published: (2025)
by: Unser, Michael, et al.
Published: (2025)
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Composite Optimization with Error Feedback: the Dual Averaging Approach
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Accelerated Distributed Optimization with Compression and Error Feedback
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Efficient Low-rank Identification via Accelerated Iteratively Reweighted Nuclear Norm Minimization
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
De-singularity Subgradient for the $q$-th-Powered $\ell_p$-Norm Weber Location Problem
by: Lai, Zhao-Rong, et al.
Published: (2024)
by: Lai, Zhao-Rong, et al.
Published: (2024)
Finite-Time Bounds for Two-Time-Scale Stochastic Approximation with Arbitrary Norm Contractions and Markovian Noise
by: Chandak, Siddharth, et al.
Published: (2025)
by: Chandak, Siddharth, et al.
Published: (2025)
Robust Least-Squares Optimization for Data-Driven Predictive Control: A Geometric Approach
by: Bharadwaj, Shreyas, et al.
Published: (2025)
by: Bharadwaj, Shreyas, et al.
Published: (2025)
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2025)
by: Li, Huan, et al.
Published: (2025)
Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence Analysis
by: Hao, Jie, et al.
Published: (2024)
by: Hao, Jie, et al.
Published: (2024)
New Lower Bounds for Stochastic Non-Convex Optimization through Divergence Decomposition
by: Saad, El Mehdi, et al.
Published: (2025)
by: Saad, El Mehdi, et al.
Published: (2025)
Lean and Mean Adaptive Optimization via Subset-Norm and Subspace-Momentum with Convergence Guarantees
by: Nguyen, Thien Hang, et al.
Published: (2024)
by: Nguyen, Thien Hang, et al.
Published: (2024)
A Second-Order Majorant Algorithm for Nonnegative Matrix Factorization
by: Pham, Mai-Quyen, et al.
Published: (2023)
by: Pham, Mai-Quyen, et al.
Published: (2023)
Iterative Regularization with k-support Norm: An Important Complement to Sparse Recovery
by: de Vazelhes, William, et al.
Published: (2023)
by: de Vazelhes, William, et al.
Published: (2023)
From Learning to Optimize to Learning Optimization Algorithms
by: Castera, Camille, et al.
Published: (2024)
by: Castera, Camille, et al.
Published: (2024)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Optimizing Posterior Samples for Bayesian Optimization via Rootfinding
by: Adebiyi, Taiwo A., et al.
Published: (2024)
by: Adebiyi, Taiwo A., et al.
Published: (2024)
Preference-Optimized Pareto Set Learning for Blackbox Optimization
by: Haishan, Zhang, et al.
Published: (2024)
by: Haishan, Zhang, et al.
Published: (2024)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Dynamic Proximal Gradient Algorithms for Schatten-$p$ Quasi-Norm Regularized Problems
by: Shen, Weiping, et al.
Published: (2026)
by: Shen, Weiping, et al.
Published: (2026)
Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data
by: Seraghiti, Giovanni, et al.
Published: (2026)
by: Seraghiti, Giovanni, et al.
Published: (2026)
Accelerating Multi-Block Constrained Optimization Through Learning to Optimize
by: Liang, Ling, et al.
Published: (2024)
by: Liang, Ling, et al.
Published: (2024)
Reinforcement learning for adaptive interior point methods in convex quadratic programming
by: Bertoncini, Jeremy, et al.
Published: (2025)
by: Bertoncini, Jeremy, et al.
Published: (2025)
Bayesian Optimization for Non-Convex Two-Stage Stochastic Optimization Problems
by: Buckingham, Jack M., et al.
Published: (2024)
by: Buckingham, Jack M., et al.
Published: (2024)
Optimizer's Information Criterion: Dissecting and Correcting Bias in Data-Driven Optimization
by: Iyengar, Garud, et al.
Published: (2023)
by: Iyengar, Garud, et al.
Published: (2023)
Similar Items
-
Modular Duality in Deep Learning
by: Bernstein, Jeremy, et al.
Published: (2024) -
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
by: Liu, Zijian
Published: (2025) -
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025) -
Small Gradient Norm Regret for Online Convex Optimization
by: Gao, Wenzhi, et al.
Published: (2026) -
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)