Grokking at the Edge of Linear Separability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beck, Alon, Levi, Noam, Bar-Sinai, Yohai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913949099753472
author Beck, Alon
Levi, Noam
Bar-Sinai, Yohai
author_facet Beck, Alon
Levi, Noam
Bar-Sinai, Yohai
contents We investigate the phenomenon of grokking -- delayed generalization accompanied by non-monotonic test loss behavior -- in a simple binary logistic classification task, for which "memorizing" and "generalizing" solutions can be strictly defined. Surprisingly, we find that grokking arises naturally even in this minimal model when the parameters of the problem are close to a critical point, and provide both empirical and analytical insights into its mechanism. Concretely, by appealing to the implicit bias of gradient descent, we show that logistic regression can exhibit grokking when the training dataset is nearly linearly separable from the origin and there is strong noise in the perpendicular directions. The underlying reason is that near the critical point, "flat" directions in the loss landscape with nearly zero gradient cause training dynamics to linger for arbitrarily long times near quasi-stable solutions before eventually reaching the global minimum. Finally, we highlight similarities between our findings and the recent literature, strengthening the conjecture that grokking generally occurs in proximity to the interpolation threshold, reminiscent of critical phenomena often observed in physical systems.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04489
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grokking at the Edge of Linear Separability
Beck, Alon
Levi, Noam
Bar-Sinai, Yohai
Machine Learning
Disordered Systems and Neural Networks
Mathematical Physics
We investigate the phenomenon of grokking -- delayed generalization accompanied by non-monotonic test loss behavior -- in a simple binary logistic classification task, for which "memorizing" and "generalizing" solutions can be strictly defined. Surprisingly, we find that grokking arises naturally even in this minimal model when the parameters of the problem are close to a critical point, and provide both empirical and analytical insights into its mechanism. Concretely, by appealing to the implicit bias of gradient descent, we show that logistic regression can exhibit grokking when the training dataset is nearly linearly separable from the origin and there is strong noise in the perpendicular directions. The underlying reason is that near the critical point, "flat" directions in the loss landscape with nearly zero gradient cause training dynamics to linger for arbitrarily long times near quasi-stable solutions before eventually reaching the global minimum. Finally, we highlight similarities between our findings and the recent literature, strengthening the conjecture that grokking generally occurs in proximity to the interpolation threshold, reminiscent of critical phenomena often observed in physical systems.
title Grokking at the Edge of Linear Separability
topic Machine Learning
Disordered Systems and Neural Networks
Mathematical Physics
url https://arxiv.org/abs/2410.04489