SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
Fuente:
arXiv
Saved in:
| Main Authors: | Labarrière, Hippolyte, Molinari, Cesare, Villa, Silvia, Rosasco, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimization Insights into Deep Diagonal Linear Networks
by: Labarrière, Hippolyte, et al.
Published: (2024)
by: Labarrière, Hippolyte, et al.
Published: (2024)
Stochastic Zeroth order Descent with Structured Directions
by: Rando, Marco, et al.
Published: (2022)
by: Rando, Marco, et al.
Published: (2022)
Convergence of zeroth-order proximal point algorithms in the high-temperature regime
by: Naldi, Emanuele, et al.
Published: (2026)
by: Naldi, Emanuele, et al.
Published: (2026)
Iterative regularization in classification via hinge loss diagonal descent
by: Apidopoulos, Vassilis, et al.
Published: (2022)
by: Apidopoulos, Vassilis, et al.
Published: (2022)
A Structured Tour of Optimization with Finite Differences
by: Rando, Marco, et al.
Published: (2025)
by: Rando, Marco, et al.
Published: (2025)
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
by: Fazla, Arda, et al.
Published: (2026)
by: Fazla, Arda, et al.
Published: (2026)
Linear quadratic control of nonlinear systems with Koopman operator learning and the Nyström method
by: Caldarelli, Edoardo, et al.
Published: (2024)
by: Caldarelli, Edoardo, et al.
Published: (2024)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
A Structured Proximal Stochastic Variance Reduced Zeroth-order Algorithm
by: Rando, Marco, et al.
Published: (2025)
by: Rando, Marco, et al.
Published: (2025)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
Variance reduction techniques for stochastic proximal point algorithms
by: Traoré, Cheik, et al.
Published: (2023)
by: Traoré, Cheik, et al.
Published: (2023)
Proximal basin hopping: global optimization with guarantees
by: Lauga, Guillaume, et al.
Published: (2026)
by: Lauga, Guillaume, et al.
Published: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Stochastic Variance-Reduced Newton: Accelerating Finite-Sum Minimization with Large Batches
by: Dereziński, Michał
Published: (2022)
by: Dereziński, Michał
Published: (2022)
Strong Convergence of FISTA Iterates under H{ö}lderian and Quadratic Growth Conditions
by: Aujol, Jean-François, et al.
Published: (2024)
by: Aujol, Jean-François, et al.
Published: (2024)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Dynamic robotic cloth folding with efficient Koopman operator-based model predictive control
by: Caldarelli, Edoardo, et al.
Published: (2026)
by: Caldarelli, Edoardo, et al.
Published: (2026)
Solving Stochastic Variational Inequalities without the Bounded Variance Assumption
by: Alacaoglu, Ahmet, et al.
Published: (2026)
by: Alacaoglu, Ahmet, et al.
Published: (2026)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Breaking the Stochasticity Barrier: An Adaptive Variance-Reduced Method for Variational Inequalities
by: Jeong, Yungi, et al.
Published: (2026)
by: Jeong, Yungi, et al.
Published: (2026)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Making SGD Parameter-Free
by: Carmon, Yair, et al.
Published: (2022)
by: Carmon, Yair, et al.
Published: (2022)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
Stochastic Gradient Langevin Dynamics with Variance Reduction
by: Huang, Zhishen, et al.
Published: (2021)
by: Huang, Zhishen, et al.
Published: (2021)
SketchySGD: Reliable Stochastic Optimization via Randomized Curvature Estimates
by: Frangella, Zachary, et al.
Published: (2022)
by: Frangella, Zachary, et al.
Published: (2022)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Online Convex Optimization with Unbounded Memory
by: Kumar, Raunak, et al.
Published: (2022)
by: Kumar, Raunak, et al.
Published: (2022)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
Demystifying SGD with Doubly Stochastic Gradients
by: Kim, Kyurae, et al.
Published: (2024)
by: Kim, Kyurae, et al.
Published: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD
by: Hu, Jie, et al.
Published: (2024)
by: Hu, Jie, et al.
Published: (2024)
Model Consistency of the Iterative Regularization of Dual Ascent for Low-Complexity Regularization
by: Gao, Jie, et al.
Published: (2025)
by: Gao, Jie, et al.
Published: (2025)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
SGD with memory: fundamental properties and stochastic acceleration
by: Yarotsky, Dmitry, et al.
Published: (2024)
by: Yarotsky, Dmitry, et al.
Published: (2024)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Similar Items
-
Optimization Insights into Deep Diagonal Linear Networks
by: Labarrière, Hippolyte, et al.
Published: (2024) -
Stochastic Zeroth order Descent with Structured Directions
by: Rando, Marco, et al.
Published: (2022) -
Convergence of zeroth-order proximal point algorithms in the high-temperature regime
by: Naldi, Emanuele, et al.
Published: (2026) -
Iterative regularization in classification via hinge loss diagonal descent
by: Apidopoulos, Vassilis, et al.
Published: (2022) -
A Structured Tour of Optimization with Finite Differences
by: Rando, Marco, et al.
Published: (2025)