A Hessian-Aware Stochastic Differential Equation for Modelling SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiang, Shen, Zebang, Zhang, Liang, He, Niao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Schrödinger Eigenfunction Method for Long-Horizon Stochastic Optimal Control
by: Claeys, Louis, et al.
Published: (2026)
by: Claeys, Louis, et al.
Published: (2026)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
Landing with the Score: Riemannian Optimization through Denoising
by: Kharitenko, Andrey, et al.
Published: (2025)
by: Kharitenko, Andrey, et al.
Published: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024)
by: Sun, Ying, et al.
Published: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
Biased Stochastic First-Order Methods for Conditional Stochastic Optimization and Applications in Meta Learning
by: Hu, Yifan, et al.
Published: (2020)
by: Hu, Yifan, et al.
Published: (2020)
Primal Methods for Variational Inequality Problems with Functional Constraints
by: Zhang, Liang, et al.
Published: (2024)
by: Zhang, Liang, et al.
Published: (2024)
On the Connectedness of Sublevel Sets in Invex Optimization
by: Thoma, Vinzenz, et al.
Published: (2026)
by: Thoma, Vinzenz, et al.
Published: (2026)
Multi-level Monte-Carlo Gradient Methods for Stochastic Optimization with Biased Oracles
by: Hu, Yifan, et al.
Published: (2024)
by: Hu, Yifan, et al.
Published: (2024)
Stochastic Hessian Fittings with Lie Groups
by: Li, Xi-Lin
Published: (2024)
by: Li, Xi-Lin
Published: (2024)
Optimal Guarantees for Algorithmic Reproducibility and Gradient Complexity in Convex Optimization
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
Demystifying SGD with Doubly Stochastic Gradients
by: Kim, Kyurae, et al.
Published: (2024)
by: Kim, Kyurae, et al.
Published: (2024)
Optimal Local Convergence Rates of Stochastic First-Order Methods under Local $α$-PL
by: Masiha, Saeed, et al.
Published: (2024)
by: Masiha, Saeed, et al.
Published: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Superquantile-Gibbs Relaxation for Minima-selection in Bilevel Optimization
by: Masiha, Saeed, et al.
Published: (2025)
by: Masiha, Saeed, et al.
Published: (2025)
On the Crucial Role of Initialization for Matrix Factorization
by: Li, Bingcong, et al.
Published: (2024)
by: Li, Bingcong, et al.
Published: (2024)
On the Benefits of Weight Normalization for Overparameterized Matrix Sensing
by: Wei, Yudong, et al.
Published: (2025)
by: Wei, Yudong, et al.
Published: (2025)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
by: Li, Chris Junchi
Published: (2024)
by: Li, Chris Junchi
Published: (2024)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
by: Cayci, Semih, et al.
Published: (2021)
by: Cayci, Semih, et al.
Published: (2021)
DPZero: Private Fine-Tuning of Language Models without Backpropagation
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
SketchySGD: Reliable Stochastic Optimization via Randomized Curvature Estimates
by: Frangella, Zachary, et al.
Published: (2022)
by: Frangella, Zachary, et al.
Published: (2022)
Select-then-differentiate: Solving Bilevel Optimization with Manifold Lower-level Solution Sets
by: Masiha, Saeed, et al.
Published: (2026)
by: Masiha, Saeed, et al.
Published: (2026)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)
by: Khah, Saleh Vatan, et al.
Published: (2025)
Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes
by: Huang, Yan, et al.
Published: (2024)
by: Huang, Yan, et al.
Published: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
PoLAR: Polar-Decomposed Low-Rank Adapter Representation
by: Lion, Kai, et al.
Published: (2025)
by: Lion, Kai, et al.
Published: (2025)
Diffusion-Based Stochastic Operator Networks for Uncertainty Quantification in Stochastic Partial Differential Equations
by: Huynh, Phuoc-Toan, et al.
Published: (2026)
by: Huynh, Phuoc-Toan, et al.
Published: (2026)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
by: Surjanovic, Nikola, et al.
Published: (2025)
by: Surjanovic, Nikola, et al.
Published: (2025)
Non-Parametric Learning of Stochastic Differential Equations with Non-asymptotic Fast Rates of Convergence
by: Bonalli, Riccardo, et al.
Published: (2023)
by: Bonalli, Riccardo, et al.
Published: (2023)
Zeroth-Order Optimization at the Edge of Stability
by: Song, Minhak, et al.
Published: (2026)
by: Song, Minhak, et al.
Published: (2026)
Achieving ${O}(ε^{-1.5})$ Complexity in Hessian/Jacobian-free Stochastic Bilevel Optimization
by: Yang, Yifan, et al.
Published: (2023)
by: Yang, Yifan, et al.
Published: (2023)
Faster Convergence of Local SGD for Over-Parameterized Models
by: Qin, Tiancheng, et al.
Published: (2022)
by: Qin, Tiancheng, et al.
Published: (2022)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Exploiting Approximate Symmetry for Efficient Multi-Agent Reinforcement Learning
by: Yardim, Batuhan, et al.
Published: (2024)
by: Yardim, Batuhan, et al.
Published: (2024)
Similar Items
-
A Schrödinger Eigenfunction Method for Long-Horizon Stochastic Optimal Control
by: Claeys, Louis, et al.
Published: (2026) -
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024) -
Landing with the Score: Riemannian Optimization through Denoising
by: Kharitenko, Andrey, et al.
Published: (2025) -
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024) -
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)