Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Mingze, Min, Zeping, Wu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
How Transformers Get Rich: Approximation and Dynamics Analysis
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024)
by: Cai, Yuhang, et al.
Published: (2024)
Fast and Effective Computation of Generalized Symmetric Matrix Factorization
by: Yang, Lei, et al.
Published: (2026)
by: Yang, Lei, et al.
Published: (2026)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Efficient Low-rank Identification via Accelerated Iteratively Reweighted Nuclear Norm Minimization
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Improving Generalization and Convergence by Enhancing Implicit Regularization
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Efficient Online Large-Margin Classification via Dual Certificates
by: Ho-Nguyen, Nam, et al.
Published: (2025)
by: Ho-Nguyen, Nam, et al.
Published: (2025)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
Fast Nonlinear Two-Time-Scale Stochastic Approximation: Achieving $O(1/k)$ Finite-Sample Complexity
by: Doan, Thinh T.
Published: (2024)
by: Doan, Thinh T.
Published: (2024)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
Old Optimizer, New Norm: An Anthology
by: Bernstein, Jeremy, et al.
Published: (2024)
by: Bernstein, Jeremy, et al.
Published: (2024)
Exponential Concentration in Stochastic Approximation
by: Law, Kody, et al.
Published: (2022)
by: Law, Kody, et al.
Published: (2022)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Optimal Transport with Tempered Exponential Measures
by: Amid, Ehsan, et al.
Published: (2023)
by: Amid, Ehsan, et al.
Published: (2023)
De-singularity Subgradient for the $q$-th-Powered $\ell_p$-Norm Weber Location Problem
by: Lai, Zhao-Rong, et al.
Published: (2024)
by: Lai, Zhao-Rong, et al.
Published: (2024)
Universal Architectures for the Learning of Polyhedral Norms and Convex Regularizers
by: Unser, Michael, et al.
Published: (2025)
by: Unser, Michael, et al.
Published: (2025)
Small Gradient Norm Regret for Online Convex Optimization
by: Gao, Wenzhi, et al.
Published: (2026)
by: Gao, Wenzhi, et al.
Published: (2026)
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Maximizing Reliability with Bayesian Optimization
by: Buckingham, Jack M., et al.
Published: (2026)
by: Buckingham, Jack M., et al.
Published: (2026)
Subsampled Ensemble Can Improve Generalization Tail Exponentially
by: Qian, Huajie, et al.
Published: (2024)
by: Qian, Huajie, et al.
Published: (2024)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Fast sparse optimization via adaptive shrinkage
by: Cerone, Vito, et al.
Published: (2025)
by: Cerone, Vito, et al.
Published: (2025)
Achieving Better Local Regret Bound for Online Non-Convex Bilevel Optimization
by: Jia, Tingkai, et al.
Published: (2026)
by: Jia, Tingkai, et al.
Published: (2026)
Robust and Fast Training via Per-Sample Clipping
by: Nobile, Davide, et al.
Published: (2026)
by: Nobile, Davide, et al.
Published: (2026)
Improving Feasibility via Fast Autoencoder-Based Projections
by: Chzhen, Maria, et al.
Published: (2026)
by: Chzhen, Maria, et al.
Published: (2026)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
by: Kale, Sacchit, et al.
Published: (2026)
by: Kale, Sacchit, et al.
Published: (2026)
Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise
by: Zhang, Jiayu, et al.
Published: (2026)
by: Zhang, Jiayu, et al.
Published: (2026)
Convex Relaxation for Solving Large-Margin Classifiers in Hyperbolic Space
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
Achieving Linear Speedup for Composite Federated Learning
by: Huang, Kun, et al.
Published: (2026)
by: Huang, Kun, et al.
Published: (2026)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
by: Kravatskiy, Alexey, et al.
Published: (2025)
by: Kravatskiy, Alexey, et al.
Published: (2025)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
(Un)supervised Learning of Maximal Lyapunov Functions
by: Barreau, Matthieu, et al.
Published: (2024)
by: Barreau, Matthieu, et al.
Published: (2024)
Active Learning For Contextual Linear Optimization: A Margin-Based Approach
by: Liu, Mo, et al.
Published: (2023)
by: Liu, Mo, et al.
Published: (2023)
Similar Items
-
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026) -
How Transformers Get Rich: Approximation and Dynamics Analysis
by: Wang, Mingze, et al.
Published: (2024) -
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024) -
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024) -
Fast and Effective Computation of Generalized Symmetric Matrix Factorization
by: Yang, Lei, et al.
Published: (2026)