Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
Fuente:
arXiv
Saved in:
| Main Authors: | Medyakov, Daniil, Molodtsov, Gleb, Chezhegov, Savelii, Rebrikov, Alexey, Beznosikov, Aleksandr |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effective Method with Compression for Distributed and Federated Cocoercive Variational Inequalities
by: Medyakov, Daniil, et al.
Published: (2024)
by: Medyakov, Daniil, et al.
Published: (2024)
Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Optimal Data Splitting in Distributed Optimization for Machine Learning
by: Medyakov, Daniil, et al.
Published: (2024)
by: Medyakov, Daniil, et al.
Published: (2024)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)
by: Bylinkin, Dmitry, et al.
Published: (2025)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Accelerated Stochastic Gradient Method with Applications to Consensus Problem in Markov-Varying Networks
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Local SGD for Near-Quadratic Problems: Improving Convergence under Unconstrained Noise Conditions
by: Sadchikov, Andrey, et al.
Published: (2024)
by: Sadchikov, Andrey, et al.
Published: (2024)
Incorporating Preconditioning into Accelerated Approaches: Theoretical Guarantees and Practical Improvement
by: Trifonov, Stepan, et al.
Published: (2025)
by: Trifonov, Stepan, et al.
Published: (2025)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
Bant: Byzantine Antidote via Trial Function and Trust Scores
by: Molodtsov, Gleb, et al.
Published: (2025)
by: Molodtsov, Gleb, et al.
Published: (2025)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Local Methods with Adaptivity via Scaling
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Unified Theory of Adaptive Variance Reduction
by: Shestakov, Aleksandr, et al.
Published: (2025)
by: Shestakov, Aleksandr, et al.
Published: (2025)
New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
by: Veprikov, Andrey, et al.
Published: (2024)
by: Veprikov, Andrey, et al.
Published: (2024)
Decentralized Finite-Sum Optimization over Time-Varying Networks
by: Metelev, Dmitry, et al.
Published: (2024)
by: Metelev, Dmitry, et al.
Published: (2024)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)
by: Khah, Saleh Vatan, et al.
Published: (2025)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
by: Prokhorov, Boris, et al.
Published: (2026)
by: Prokhorov, Boris, et al.
Published: (2026)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
First Order Methods with Markovian Noise: from Acceleration to Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Improved Last-Iterate Convergence of Shuffling Gradient Methods for Nonsmooth Convex Optimization
by: Liu, Zijian, et al.
Published: (2025)
by: Liu, Zijian, et al.
Published: (2025)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
by: Ustimenko, Aleksei, et al.
Published: (2023)
by: Ustimenko, Aleksei, et al.
Published: (2023)
On the Last-Iterate Convergence of Shuffling Gradient Methods
by: Liu, Zijian, et al.
Published: (2024)
by: Liu, Zijian, et al.
Published: (2024)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Stochastic Gradient Langevin Dynamics with Variance Reduction
by: Huang, Zhishen, et al.
Published: (2021)
by: Huang, Zhishen, et al.
Published: (2021)
Shuffling Gradient-Based Methods for Nonconvex-Concave Minimax Optimization
by: Tran-Dinh, Quoc, et al.
Published: (2024)
by: Tran-Dinh, Quoc, et al.
Published: (2024)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Gradient Estimation and Variance Reduction in Stochastic and Deterministic Models
by: Keane, Ronan
Published: (2024)
by: Keane, Ronan
Published: (2024)
Activations and Gradients Compression for Model-Parallel Training
by: Rudakov, Mikhail, et al.
Published: (2024)
by: Rudakov, Mikhail, et al.
Published: (2024)
Communication-Efficient Federated Learning with Adaptive Number of Participants
by: Skorik, Sergey, et al.
Published: (2025)
by: Skorik, Sergey, et al.
Published: (2025)
Methods for Optimization Problems with Markovian Stochasticity and Non-Euclidean Geometry
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
Shuffling Momentum Gradient Algorithm for Convex Optimization
by: Tran, Trang H., et al.
Published: (2024)
by: Tran, Trang H., et al.
Published: (2024)
Projected Forward Gradient-Guided Frank-Wolfe Algorithm via Variance Reduction
by: Rostami, M., et al.
Published: (2024)
by: Rostami, M., et al.
Published: (2024)
Global Convergence of Natural Policy Gradient with Hessian-aided Momentum Variance Reduction
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Methods for Solving Variational Inequalities with Markovian Stochasticity
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Similar Items
-
Effective Method with Compression for Distributed and Federated Cocoercive Variational Inequalities
by: Medyakov, Daniil, et al.
Published: (2024) -
Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
by: Medyakov, Daniil, et al.
Published: (2025) -
Optimal Data Splitting in Distributed Optimization for Machine Learning
by: Medyakov, Daniil, et al.
Published: (2024) -
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025) -
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)