Why Do We Need Warm-up? A Theoretical Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Alimisis, Foivos, Islamov, Rustem, Lucchi, Aurelien |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Safe-EF: Error Feedback for Nonsmooth Constrained Optimization
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Characterization of optimization problems that are solvable iteratively with linear convergence
by: Alimisis, Foivos
Published: (2024)
by: Alimisis, Foivos
Published: (2024)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Towards Faster Decentralized Stochastic Optimization with Communication Compression
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
A geodesic convexity-like structure for the polar decomposition of a square matrix
by: Alimisis, Foivos, et al.
Published: (2024)
by: Alimisis, Foivos, et al.
Published: (2024)
A Nesterov-style Accelerated Gradient Descent Algorithm for the Symmetric Eigenvalue Problem
by: Alimisis, Foivos, et al.
Published: (2024)
by: Alimisis, Foivos, et al.
Published: (2024)
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
SDEs for Minimax Optimization
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective
by: Compagnoni, Enea Monzio, et al.
Published: (2026)
by: Compagnoni, Enea Monzio, et al.
Published: (2026)
Gradient-type subspace iteration methods for the symmetric eigenvalue problem
by: Alimisis, Foivos, et al.
Published: (2023)
by: Alimisis, Foivos, et al.
Published: (2023)
Regret-Optimal Federated Transfer Learning for Kernel Regression with Applications in American Option Pricing
by: Yang, Xuwei, et al.
Published: (2023)
by: Yang, Xuwei, et al.
Published: (2023)
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
by: Lin, Wu, et al.
Published: (2024)
by: Lin, Wu, et al.
Published: (2024)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Sporadic Gradient Tracking over Directed Graphs: A Theoretical Perspective on Decentralized Federated Learning
by: Zehtabi, Shahryar, et al.
Published: (2026)
by: Zehtabi, Shahryar, et al.
Published: (2026)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Controllable Expensive Multi-objective Learning with Warm-starting Bayesian Optimization
by: Nguyen, Quang-Huy, et al.
Published: (2023)
by: Nguyen, Quang-Huy, et al.
Published: (2023)
Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching
by: Angell, Rico, et al.
Published: (2023)
by: Angell, Rico, et al.
Published: (2023)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Deconfounded Warm-Start Thompson Sampling with Applications to Precision Medicine
by: Jaiswal, Prateek, et al.
Published: (2025)
by: Jaiswal, Prateek, et al.
Published: (2025)
A Systems-Theoretic View on the Convergence of Algorithms under Disturbances
by: Er, Guner Dilsad, et al.
Published: (2025)
by: Er, Guner Dilsad, et al.
Published: (2025)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
by: Schaipp, Fabian, et al.
Published: (2026)
by: Schaipp, Fabian, et al.
Published: (2026)
Error whitening: Why Gauss-Newton outperforms Newton
by: McKay, Maricela Best, et al.
Published: (2026)
by: McKay, Maricela Best, et al.
Published: (2026)
Deterministic Global Optimization of the Acquisition Function in Bayesian Optimization: To Do or Not To Do?
by: Georgiou, Anastasia, et al.
Published: (2025)
by: Georgiou, Anastasia, et al.
Published: (2025)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm
by: Kumari, Sakshi, et al.
Published: (2026)
by: Kumari, Sakshi, et al.
Published: (2026)
Why Smooth Stability Assumptions Fail for ReLU Learning
by: Katende, Ronald
Published: (2025)
by: Katende, Ronald
Published: (2025)
Towards Robust Learning to Optimize with Theoretical Guarantees
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
A Control Theoretic Framework for Adaptive Gradient Optimizers in Machine Learning
by: Chakrabarti, Kushal, et al.
Published: (2022)
by: Chakrabarti, Kushal, et al.
Published: (2022)
Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation
by: Sokolov, Igor, et al.
Published: (2025)
by: Sokolov, Igor, et al.
Published: (2025)
Control Theoretic Approach to Fine-Tuning and Transfer Learning
by: Bayram, Erkan, et al.
Published: (2024)
by: Bayram, Erkan, et al.
Published: (2024)
Theoretical Analysis of Heteroscedastic Gaussian Processes with Posterior Distributions
by: Ito, Yuji
Published: (2024)
by: Ito, Yuji
Published: (2024)
Reevaluating Theoretical Analysis Methods for Optimization in Deep Learning
by: Tran, Hoang, et al.
Published: (2024)
by: Tran, Hoang, et al.
Published: (2024)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
Similar Items
-
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025) -
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024) -
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025) -
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026) -
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)