Grokking vs. Learning: Same Features, Different Encodings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Manning-Coe, Dmitry, Gliozzi, Jacopo, Stapleton, Alexander G., Hirst, Edward, De Tomasi, Giuseppe, Bradlyn, Barry, Berman, David S.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912217654362112
author Manning-Coe, Dmitry
Gliozzi, Jacopo
Stapleton, Alexander G.
Hirst, Edward
De Tomasi, Giuseppe
Bradlyn, Barry
Berman, David S.
author_facet Manning-Coe, Dmitry
Gliozzi, Jacopo
Stapleton, Alexander G.
Hirst, Edward
De Tomasi, Giuseppe
Bradlyn, Barry
Berman, David S.
contents Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental differences in the learned models. To do so we compare the features, compressibility, and learning dynamics of models trained via each path in two tasks. We find that grokked and steadily trained models learn the same features, but there can be large differences in the efficiency with which these features are encoded. In particular, we find a novel "compressive regime" of steady training in which there emerges a linear trade-off between model loss and compressibility, and which is absent in grokking. In this regime, we can achieve compression factors 25x times the base model, and 5x times the compression achieved in grokking. We then track how model features and compressibility develop through training. We show that model development in grokking is task-dependent, and that peak compressibility is achieved immediately after the grokking plateau. Finally, novel information-geometric measures are introduced which demonstrate that models undergoing grokking follow a straight path in information space.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01739
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grokking vs. Learning: Same Features, Different Encodings
Manning-Coe, Dmitry
Gliozzi, Jacopo
Stapleton, Alexander G.
Hirst, Edward
De Tomasi, Giuseppe
Bradlyn, Barry
Berman, David S.
Machine Learning
Disordered Systems and Neural Networks
Artificial Intelligence
Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental differences in the learned models. To do so we compare the features, compressibility, and learning dynamics of models trained via each path in two tasks. We find that grokked and steadily trained models learn the same features, but there can be large differences in the efficiency with which these features are encoded. In particular, we find a novel "compressive regime" of steady training in which there emerges a linear trade-off between model loss and compressibility, and which is absent in grokking. In this regime, we can achieve compression factors 25x times the base model, and 5x times the compression achieved in grokking. We then track how model features and compressibility develop through training. We show that model development in grokking is task-dependent, and that peak compressibility is achieved immediately after the grokking plateau. Finally, novel information-geometric measures are introduced which demonstrate that models undergoing grokking follow a straight path in information space.
title Grokking vs. Learning: Same Features, Different Encodings
topic Machine Learning
Disordered Systems and Neural Networks
Artificial Intelligence
url https://arxiv.org/abs/2502.01739