Refresh-Scaling the Memory of Balanced Adam

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fernández-Hernández, Alberto, Pérez-Corral, Cristian, Mestre, Jose I., Dolz, Manuel F., Quintana-Ortí, Enrique S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909035024875520
author Fernández-Hernández, Alberto
Pérez-Corral, Cristian
Mestre, Jose I.
Dolz, Manuel F.
Quintana-Ortí, Enrique S.
author_facet Fernández-Hernández, Alberto
Pérez-Corral, Cristian
Mestre, Jose I.
Dolz, Manuel F.
Quintana-Ortí, Enrique S.
contents Recent evidence suggests that Adam performs robustly when its momentum parameters are tied, $β_1=β_2$, reducing the optimizer to a single remaining parameter. However, how this parameter should be set remains poorly understood. We argue that, in balanced Adam, $β$ should not be treated as a dimensionless constant: it defines a statistical memory horizon $H_β=(1-β)^{-1}$. In terms of the effective learning horizon $T_{\mathrm{ES}}$, estimated from the validation trajectory, we study the refresh count $R_β=(1-β)T_{\mathrm{ES}}$, which measures how many times Adam renews its internal statistics during the useful phase of training. Across 11 vision and language experiments, we find that choosing $β$ so that $R_β\approx1000$ selects different $β$ values depending on the training scale, yet improves robustness over the best fixed-beta baseline. Compared with the strongest fixed choice $β=0.944$, the refresh rule improves worst-case robustness, reducing the maximum relative gap in validation loss by 33.4\%, while bringing all 11 runs within 1\% of their validation oracle. These results suggest that the remaining hyperparameter of balanced Adam is more naturally viewed as a memory-scale variable than as a fixed constant. This provides a simple budget-aware perspective on optimizer scaling and opens a path toward treating Adam's momentum as part of the learning dynamics rather than as a static default.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10119
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Refresh-Scaling the Memory of Balanced Adam
Fernández-Hernández, Alberto
Pérez-Corral, Cristian
Mestre, Jose I.
Dolz, Manuel F.
Quintana-Ortí, Enrique S.
Machine Learning
Recent evidence suggests that Adam performs robustly when its momentum parameters are tied, $β_1=β_2$, reducing the optimizer to a single remaining parameter. However, how this parameter should be set remains poorly understood. We argue that, in balanced Adam, $β$ should not be treated as a dimensionless constant: it defines a statistical memory horizon $H_β=(1-β)^{-1}$. In terms of the effective learning horizon $T_{\mathrm{ES}}$, estimated from the validation trajectory, we study the refresh count $R_β=(1-β)T_{\mathrm{ES}}$, which measures how many times Adam renews its internal statistics during the useful phase of training. Across 11 vision and language experiments, we find that choosing $β$ so that $R_β\approx1000$ selects different $β$ values depending on the training scale, yet improves robustness over the best fixed-beta baseline. Compared with the strongest fixed choice $β=0.944$, the refresh rule improves worst-case robustness, reducing the maximum relative gap in validation loss by 33.4\%, while bringing all 11 runs within 1\% of their validation oracle. These results suggest that the remaining hyperparameter of balanced Adam is more naturally viewed as a memory-scale variable than as a fixed constant. This provides a simple budget-aware perspective on optimizer scaling and opens a path toward treating Adam's momentum as part of the learning dynamics rather than as a static default.
title Refresh-Scaling the Memory of Balanced Adam
topic Machine Learning
url https://arxiv.org/abs/2605.10119