Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
Fuente:
arXiv
Saved in:
| Main Authors: | Riabinin, Artem, Veprikov, Andrey, Bolatov, Arman, Takáč, Martin, Beznosikov, Aleksandr |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
WeightLoRA: Keep Only Necessary Adapters
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026)
by: Parfenov, Valery, et al.
Published: (2026)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
by: Veprikov, Andrey, et al.
Published: (2024)
by: Veprikov, Andrey, et al.
Published: (2024)
Methods for Optimization Problems with Markovian Stochasticity and Non-Euclidean Geometry
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
A Novel Unified Parametric Assumption for Nonconvex Optimization
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
by: Ustimenko, Aleksei, et al.
Published: (2023)
by: Ustimenko, Aleksei, et al.
Published: (2023)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Exploring New Frontiers in Vertical Federated Learning: the Role of Saddle Point Reformulation
by: Beznosikov, Aleksandr, et al.
Published: (2026)
by: Beznosikov, Aleksandr, et al.
Published: (2026)
Optimal Data Splitting in Distributed Optimization for Machine Learning
by: Medyakov, Daniil, et al.
Published: (2024)
by: Medyakov, Daniil, et al.
Published: (2024)
Markovian Compression: Looking to the Past Helps Accelerate the Future
by: Veprikov, Andrey, et al.
Published: (2026)
by: Veprikov, Andrey, et al.
Published: (2026)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Local Methods with Adaptivity via Scaling
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Bant: Byzantine Antidote via Trial Function and Trust Scores
by: Molodtsov, Gleb, et al.
Published: (2025)
by: Molodtsov, Gleb, et al.
Published: (2025)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)
by: Bylinkin, Dmitry, et al.
Published: (2025)
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
by: Prokhorov, Boris, et al.
Published: (2026)
by: Prokhorov, Boris, et al.
Published: (2026)
Decentralized Personalized Federated Learning for Min-Max Problems
by: Borodich, Ekaterina, et al.
Published: (2021)
by: Borodich, Ekaterina, et al.
Published: (2021)
Adaptive Regularized Newton Method with Inexact Hessian
by: Shestakov, Aleksandr, et al.
Published: (2025)
by: Shestakov, Aleksandr, et al.
Published: (2025)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
by: Gorbunov, Eduard, et al.
Published: (2024)
by: Gorbunov, Eduard, et al.
Published: (2024)
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
by: Yudin, Nikolay, et al.
Published: (2025)
by: Yudin, Nikolay, et al.
Published: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
by: Agafonov, Artem, et al.
Published: (2025)
by: Agafonov, Artem, et al.
Published: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
First Order Methods with Markovian Noise: from Acceleration to Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
Similar Items
-
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025) -
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
by: Bolatov, Arman, et al.
Published: (2026) -
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021) -
WeightLoRA: Keep Only Necessary Adapters
by: Veprikov, Andrey, et al.
Published: (2025) -
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026)