Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Levi, Noam, Beck, Alon, Bar-Sinai, Yohai
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913222740672512
author Levi, Noam
Beck, Alon
Bar-Sinai, Yohai
author_facet Levi, Noam
Beck, Alon
Bar-Sinai, Yohai
contents Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data. We show both analytically and numerically that grokking can surprisingly occur in linear networks performing linear tasks in a simple teacher-student setup with Gaussian inputs. In this setting, the full training dynamics is derived in terms of the training and generalization data covariance matrix. We present exact predictions on how the grokking time depends on input and output dimensionality, train sample size, regularization, and network initialization. We demonstrate that the sharp increase in generalization accuracy may not imply a transition from "memorization" to "understanding", but can simply be an artifact of the accuracy measure. We provide empirical verification for our calculations, along with preliminary results indicating that some predictions also hold for deeper networks, with non-linear activations.
format Preprint
id arxiv_https___arxiv_org_abs_2310_16441
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
Levi, Noam
Beck, Alon
Bar-Sinai, Yohai
Machine Learning
Disordered Systems and Neural Networks
Mathematical Physics
Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data. We show both analytically and numerically that grokking can surprisingly occur in linear networks performing linear tasks in a simple teacher-student setup with Gaussian inputs. In this setting, the full training dynamics is derived in terms of the training and generalization data covariance matrix. We present exact predictions on how the grokking time depends on input and output dimensionality, train sample size, regularization, and network initialization. We demonstrate that the sharp increase in generalization accuracy may not imply a transition from "memorization" to "understanding", but can simply be an artifact of the accuracy measure. We provide empirical verification for our calculations, along with preliminary results indicating that some predictions also hold for deeper networks, with non-linear activations.
title Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
topic Machine Learning
Disordered Systems and Neural Networks
Mathematical Physics
url https://arxiv.org/abs/2310.16441