Cumulative Learning Rate Adaptation: Revisiting Path-Based Schedules for SGD and Adam
Fuente:
arXiv
Saved in:
| Main Authors: | Atamna, Asma, Maus, Tom, Kievelitz, Fabian, Glasmachers, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Genetic Algorithms for Efficient Demonstration Generation in Real-World Reinforcement Learning Environments
by: Maus, Tom, et al.
Published: (2025)
by: Maus, Tom, et al.
Published: (2025)
Balancing Specialization and Centralization: A Multi-Agent Reinforcement Learning Benchmark for Sequential Industrial Control
by: Maus, Tom, et al.
Published: (2025)
by: Maus, Tom, et al.
Published: (2025)
Solving a Real-World Optimization Problem Using Proximal Policy Optimization with Curriculum Learning and Reward Engineering
by: Pendyala, Abhijeet, et al.
Published: (2024)
by: Pendyala, Abhijeet, et al.
Published: (2024)
Evolutionary Warm-Starts for Reinforcement Learning in Industrial Continuous Control
by: Maus, Tom, et al.
Published: (2026)
by: Maus, Tom, et al.
Published: (2026)
SortingEnv: An Extendable RL-Environment for an Industrial Sorting Process
by: Maus, Tom, et al.
Published: (2025)
by: Maus, Tom, et al.
Published: (2025)
Deep Reinforcement Learning Based Navigation with Macro Actions and Topological Maps
by: Hakenes, Simon, et al.
Published: (2025)
by: Hakenes, Simon, et al.
Published: (2025)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026)
by: Glentis, Athanasios, et al.
Published: (2026)
Curriculum RL meets Monte Carlo Planning: Optimization of a Real World Container Management Problem
by: Pendyala, Abhijeet, et al.
Published: (2025)
by: Pendyala, Abhijeet, et al.
Published: (2025)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Tuning Learning Rates with the Cumulative-Learning Constant
by: Faraj, Nathan
Published: (2025)
by: Faraj, Nathan
Published: (2025)
Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Revisiting LocalSGD and SCAFFOLD: Improved Rates and Missing Analysis
by: Luo, Ruichen, et al.
Published: (2025)
by: Luo, Ruichen, et al.
Published: (2025)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
by: Mukherjee, Sagnik, et al.
Published: (2026)
by: Mukherjee, Sagnik, et al.
Published: (2026)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
Private and Fair Machine Learning: Revisiting the Disparate Impact of Differentially Private SGD
by: Demelius, Lea, et al.
Published: (2025)
by: Demelius, Lea, et al.
Published: (2025)
From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Deep-learning-based identification of individual motion characteristics from upper-limb trajectories towards disorder stage evaluation
by: Sziburis, Tim, et al.
Published: (2024)
by: Sziburis, Tim, et al.
Published: (2024)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Revisiting Learning Rate Control
by: Henheik, Micha, et al.
Published: (2025)
by: Henheik, Micha, et al.
Published: (2025)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Decoupled Relative Learning Rate Schedules
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
ProtoP-OD: Explainable Object Detection with Prototypical Parts
by: Rath-Manakidis, Pavlos, et al.
Published: (2024)
by: Rath-Manakidis, Pavlos, et al.
Published: (2024)
Implicit Bias in Noisy-SGD: With Applications to Differentially Private Training
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
by: Surjanovic, Nikola, et al.
Published: (2025)
by: Surjanovic, Nikola, et al.
Published: (2025)
Guess-and-Learn (G&L): Measuring the Cumulative Error Cost of Cold-Start Adaptation
by: Arnold, Roland
Published: (2025)
by: Arnold, Roland
Published: (2025)
Comparison of Outlier Detection Algorithms on String Data
by: Maus, Philip
Published: (2026)
by: Maus, Philip
Published: (2026)
Adam-mini: Use Fewer Learning Rates To Gain More
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
An Adaptive Volatility-based Learning Rate Scheduler
by: Ren, Kieran Chai Kai
Published: (2025)
by: Ren, Kieran Chai Kai
Published: (2025)
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
by: Schaipp, Fabian, et al.
Published: (2025)
by: Schaipp, Fabian, et al.
Published: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
Similar Items
-
Leveraging Genetic Algorithms for Efficient Demonstration Generation in Real-World Reinforcement Learning Environments
by: Maus, Tom, et al.
Published: (2025) -
Balancing Specialization and Centralization: A Multi-Agent Reinforcement Learning Benchmark for Sequential Industrial Control
by: Maus, Tom, et al.
Published: (2025) -
Solving a Real-World Optimization Problem Using Proximal Policy Optimization with Curriculum Learning and Reward Engineering
by: Pendyala, Abhijeet, et al.
Published: (2024) -
Evolutionary Warm-Starts for Reinforcement Learning in Industrial Continuous Control
by: Maus, Tom, et al.
Published: (2026) -
SortingEnv: An Extendable RL-Environment for an Industrial Sorting Process
by: Maus, Tom, et al.
Published: (2025)