Logarithmic-time Schedules for Scaling Language Models with Momentum

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferbach, Damien, Paquette, Courtney, Gidel, Gauthier, Everett, Katie, Paquette, Elliot
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911454989385728
author Ferbach, Damien
Paquette, Courtney
Gidel, Gauthier
Everett, Katie
Paquette, Elliot
author_facet Ferbach, Damien
Paquette, Courtney
Gidel, Gauthier
Everett, Katie
Paquette, Elliot
contents In practice, the hyperparameters $(β_1, β_2)$ and weight-decay $λ$ in AdamW are typically kept at fixed values. Is there any reason to do otherwise? We show that for large-scale language model training, the answer is yes: by exploiting the power-law structure of language data, one can design time-varying schedules for $(β_1, β_2, λ)$ that deliver substantial performance gains. We study logarithmic-time scheduling, in which the optimizer's gradient memory horizon grows with training time. Although naive variants of this are unstable, we show that suitable damping mechanisms restore stability while preserving the benefits of longer memory. Based on this, we present ADANA, an AdamW-like optimizer that couples log-time schedules with explicit damping to balance stability and performance. We empirically evaluate ADANA across transformer scalings (45M to 2.6B parameters), comparing against AdamW, Muon, and AdEMAMix. When properly tuned, ADANA achieves up to 40% compute efficiency relative to a tuned AdamW, with gains that persist--and even improve--as model scale increases. We further show that similar benefits arise when applying logarithmic-time scheduling to AdEMAMix, and that logarithmic-time weight-decay alone can yield significant improvements. Finally, we present variants of ADANA that mitigate potential failure modes and improve robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05298
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Logarithmic-time Schedules for Scaling Language Models with Momentum
Ferbach, Damien
Paquette, Courtney
Gidel, Gauthier
Everett, Katie
Paquette, Elliot
Machine Learning
Optimization and Control
In practice, the hyperparameters $(β_1, β_2)$ and weight-decay $λ$ in AdamW are typically kept at fixed values. Is there any reason to do otherwise? We show that for large-scale language model training, the answer is yes: by exploiting the power-law structure of language data, one can design time-varying schedules for $(β_1, β_2, λ)$ that deliver substantial performance gains. We study logarithmic-time scheduling, in which the optimizer's gradient memory horizon grows with training time. Although naive variants of this are unstable, we show that suitable damping mechanisms restore stability while preserving the benefits of longer memory. Based on this, we present ADANA, an AdamW-like optimizer that couples log-time schedules with explicit damping to balance stability and performance. We empirically evaluate ADANA across transformer scalings (45M to 2.6B parameters), comparing against AdamW, Muon, and AdEMAMix. When properly tuned, ADANA achieves up to 40% compute efficiency relative to a tuned AdamW, with gains that persist--and even improve--as model scale increases. We further show that similar benefits arise when applying logarithmic-time scheduling to AdEMAMix, and that logarithmic-time weight-decay alone can yield significant improvements. Finally, we present variants of ADANA that mitigate potential failure modes and improve robustness.
title Logarithmic-time Schedules for Scaling Language Models with Momentum
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2602.05298