Saved in:
Bibliographic Details
Main Authors: Muscarnera, Luca, Estévez, Silas Ruhrberg, Xiao, Yuanzhang, Van der Schaar, Mihaela
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2606.01521
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910279146668032
author Muscarnera, Luca
Estévez, Silas Ruhrberg
Xiao, Yuanzhang
Van der Schaar, Mihaela
author_facet Muscarnera, Luca
Estévez, Silas Ruhrberg
Xiao, Yuanzhang
Van der Schaar, Mihaela
contents A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples. This gap is especially acute in high-dimensional, low-sample regimes, where many interpolating solutions exist and optimization must implicitly select among minima with different generalization properties. Following recent theoretical advances on optimization dynamics near the interpolation threshold, we note that the two-regime structure of risk minimization, with loss minimization followed by complexity minimization, motivates a biphasic optimization schedule. We thus theoretically demonstrate that GROKtimizer, a biphasic strategy that combines rapid convergence to interpolation with Critically Damped Momentum (CDM)-based post-interpolation norm minimization, offers a natural solution for selecting low-norm interpolating solutions. Under a local quadratic model of the post-interpolation basin, GROKtimizer provides a quadratic speedup over classical gradient descent, with provable optimality among first-order optimizers. To showcase the applicability of our method, we evaluate GROKtimizer on several synthetic benchmarks common in the classical grokking literature and on various real-world datasets. Finally, we reconcile our findings with the flat-minima hypothesis, highlighting the importance of post-interpolation dynamics in the construction of high-quality, generalizing models.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01521
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fast Generalization after Interpolation via Critically Damped Momentum Optimization
Muscarnera, Luca
Estévez, Silas Ruhrberg
Xiao, Yuanzhang
Van der Schaar, Mihaela
Machine Learning
A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples. This gap is especially acute in high-dimensional, low-sample regimes, where many interpolating solutions exist and optimization must implicitly select among minima with different generalization properties. Following recent theoretical advances on optimization dynamics near the interpolation threshold, we note that the two-regime structure of risk minimization, with loss minimization followed by complexity minimization, motivates a biphasic optimization schedule. We thus theoretically demonstrate that GROKtimizer, a biphasic strategy that combines rapid convergence to interpolation with Critically Damped Momentum (CDM)-based post-interpolation norm minimization, offers a natural solution for selecting low-norm interpolating solutions. Under a local quadratic model of the post-interpolation basin, GROKtimizer provides a quadratic speedup over classical gradient descent, with provable optimality among first-order optimizers. To showcase the applicability of our method, we evaluate GROKtimizer on several synthetic benchmarks common in the classical grokking literature and on various real-world datasets. Finally, we reconcile our findings with the flat-minima hypothesis, highlighting the importance of post-interpolation dynamics in the construction of high-quality, generalizing models.
title Fast Generalization after Interpolation via Critically Damped Momentum Optimization
topic Machine Learning
url https://arxiv.org/abs/2606.01521