Logarithmic-time Schedules for Scaling Language Models with Momentum
Fuente:
arXiv
Saved in:
| Main Authors: | Ferbach, Damien, Paquette, Courtney, Gidel, Gauthier, Everett, Katie, Paquette, Elliot |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026)
by: Everett, Katie, et al.
Published: (2026)
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024)
by: Paquette, Elliot, et al.
Published: (2024)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
Mirror Descent Algorithms with Nearly Dimension-Independent Rates for Differentially-Private Stochastic Saddle-Point Problems
by: González, Tomás, et al.
Published: (2024)
by: González, Tomás, et al.
Published: (2024)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)
by: Kim, Junhyung Lyle, et al.
Published: (2022)
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
by: Ferbach, Damien, et al.
Published: (2023)
by: Ferbach, Damien, et al.
Published: (2023)
Solving Hidden Monotone Variational Inequalities with Surrogate Losses
by: D'Orazio, Ryan, et al.
Published: (2024)
by: D'Orazio, Ryan, et al.
Published: (2024)
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2023)
by: Gorbunov, Eduard, et al.
Published: (2023)
Random Scaling and Momentum for Non-smooth Non-convex Optimization
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Logarithmic regret bounds for continuous-time average-reward Markov decision processes
by: Gao, Xuefeng, et al.
Published: (2022)
by: Gao, Xuefeng, et al.
Published: (2022)
RanSOM: Second-Order Momentum with Randomized Scaling for Constrained and Unconstrained Optimization
by: Chayti, El Mahdi
Published: (2026)
by: Chayti, El Mahdi
Published: (2026)
DOVA-PATBM: An Intelligent, Adaptive, and Scalable Framework for Optimizing Large-Scale EV Charging Infrastructure
by: Li, Chuan, et al.
Published: (2025)
by: Li, Chuan, et al.
Published: (2025)
Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences
by: Ferbach, Damien, et al.
Published: (2024)
by: Ferbach, Damien, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)
by: Xie, Xingyu, et al.
Published: (2022)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Improving Stochastic Cubic Newton with Momentum
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Towards Practical Second-Order Optimizers in Deep Learning: Insights from Fisher Information Analysis
by: Gomes, Damien Martins
Published: (2025)
by: Gomes, Damien Martins
Published: (2025)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
Shuffling Momentum Gradient Algorithm for Convex Optimization
by: Tran, Trang H., et al.
Published: (2024)
by: Tran, Trang H., et al.
Published: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Online Nonstochastic Prediction: Logarithmic Regret via Predictive Online Least Squares
by: Pai, Chih-Fan, et al.
Published: (2026)
by: Pai, Chih-Fan, et al.
Published: (2026)
Adaptive Optimization via Momentum on Variance-Normalized Gradients
by: Patitucci, Francisco, et al.
Published: (2026)
by: Patitucci, Francisco, et al.
Published: (2026)
Adaptive Momentum and Nonlinear Damping for Neural Network Training
by: Karoni, Aikaterini, et al.
Published: (2026)
by: Karoni, Aikaterini, et al.
Published: (2026)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Non-convex Stochastic Composite Optimization with Polyak Momentum
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)
by: Dang, Anh, et al.
Published: (2024)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
Improved Analysis for Sign-based Methods with Momentum Updates
by: Jiang, Wei, et al.
Published: (2025)
by: Jiang, Wei, et al.
Published: (2025)
Logarithmic Regret for Unconstrained Submodular Maximization Stochastic Bandit
by: Zhou, Julien, et al.
Published: (2024)
by: Zhou, Julien, et al.
Published: (2024)
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
Stochastic Compositional Optimization via Hybrid Momentum Frank--Wolfe
by: Chayti, El Mahdi
Published: (2026)
by: Chayti, El Mahdi
Published: (2026)
Better LMO-based Momentum Methods with Second-Order Information
by: Khirirat, Sarit, et al.
Published: (2025)
by: Khirirat, Sarit, et al.
Published: (2025)
Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex Optimization
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Similar Items
-
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025) -
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026) -
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024) -
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026) -
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)