Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiayu, Lin, Tianyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Stochastic Weakly Convex Optimization Under Heavy-Tailed Noises
by: Zhu, Tianxi, et al.
Published: (2025)
by: Zhu, Tianxi, et al.
Published: (2025)
Near-Optimal Decentralized Stochastic Nonconvex Optimization with Heavy-Tailed Noise
by: Wang, Menglian, et al.
Published: (2026)
by: Wang, Menglian, et al.
Published: (2026)
Optimal Asynchronous Stochastic Nonconvex Optimization under Heavy-Tailed Noise
by: Wu, Yidong, et al.
Published: (2026)
by: Wu, Yidong, et al.
Published: (2026)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
High Probability Complexity Bounds for Non-Smooth Stochastic Optimization with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2021)
by: Gorbunov, Eduard, et al.
Published: (2021)
Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping
by: Liu, Zijian, et al.
Published: (2024)
by: Liu, Zijian, et al.
Published: (2024)
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis
by: Liu, Zijian
Published: (2025)
by: Liu, Zijian
Published: (2025)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
In-Expectation Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
Concentration of General Stochastic Approximation Under Heavy-Tailed Markovian Noise
by: Agrawal, Shubhada, et al.
Published: (2026)
by: Agrawal, Shubhada, et al.
Published: (2026)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2023)
by: Gorbunov, Eduard, et al.
Published: (2023)
Heavy-Tailed and Long-Range Dependent Noise in Stochastic Approximation: A Finite-Time Analysis
by: Chandak, Siddharth, et al.
Published: (2026)
by: Chandak, Siddharth, et al.
Published: (2026)
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
by: Dang, Thanh, et al.
Published: (2025)
by: Dang, Thanh, et al.
Published: (2025)
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness
by: Kornilov, Nikita, et al.
Published: (2025)
by: Kornilov, Nikita, et al.
Published: (2025)
Breaking the Heavy-Tailed Noise Barrier in Stochastic Optimization Problems
by: Puchkin, Nikita, et al.
Published: (2023)
by: Puchkin, Nikita, et al.
Published: (2023)
Boosting-Enabled Robust System Identification of Partially Observed LTI Systems Under Heavy-Tailed Noise
by: Kanakeri, Vinay, et al.
Published: (2025)
by: Kanakeri, Vinay, et al.
Published: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
by: Liu, Zijian
Published: (2025)
by: Liu, Zijian
Published: (2025)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Neural Combinatorial Optimization with Heavy Decoder: Toward Large Scale Generalization
by: Luo, Fu, et al.
Published: (2023)
by: Luo, Fu, et al.
Published: (2023)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
Finite-Time Bounds for Two-Time-Scale Stochastic Approximation with Arbitrary Norm Contractions and Markovian Noise
by: Chandak, Siddharth, et al.
Published: (2025)
by: Chandak, Siddharth, et al.
Published: (2025)
Efficient Private SCO for Heavy-Tailed Data via Averaged Clipping
by: Jin, Chenhan, et al.
Published: (2022)
by: Jin, Chenhan, et al.
Published: (2022)
Old Optimizer, New Norm: An Anthology
by: Bernstein, Jeremy, et al.
Published: (2024)
by: Bernstein, Jeremy, et al.
Published: (2024)
Geometry of Critical Sets and Existence of Saddle Branches for Two-layer Neural Networks
by: Zhang, Leyang, et al.
Published: (2024)
by: Zhang, Leyang, et al.
Published: (2024)
Towards Guided Descent: Optimization Algorithms for Training Neural Networks At Scale
by: Nagwekar, Ansh
Published: (2025)
by: Nagwekar, Ansh
Published: (2025)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)
by: Dang, Anh, et al.
Published: (2024)
Distributed Stochastic Optimization under Heavy-Tailed Noises
by: Sun, Chao, et al.
Published: (2023)
by: Sun, Chao, et al.
Published: (2023)
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023)
by: Pan, Rui, et al.
Published: (2023)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
Similar Items
-
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024) -
Stochastic Weakly Convex Optimization Under Heavy-Tailed Noises
by: Zhu, Tianxi, et al.
Published: (2025) -
Near-Optimal Decentralized Stochastic Nonconvex Optimization with Heavy-Tailed Noise
by: Wang, Menglian, et al.
Published: (2026) -
Optimal Asynchronous Stochastic Nonconvex Optimization under Heavy-Tailed Noise
by: Wu, Yidong, et al.
Published: (2026) -
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)