Error dynamics of mini-batch gradient descent with random reshuffling for least squares regression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lok, Jackie, Sonthalia, Rishi, Rebrova, Elizaveta
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909473961934848
author Lok, Jackie
Sonthalia, Rishi
Rebrova, Elizaveta
author_facet Lok, Jackie
Sonthalia, Rishi
Rebrova, Elizaveta
contents We study the discrete dynamics of mini-batch gradient descent with random reshuffling for least squares regression. We show that the training and generalization errors depend on a sample cross-covariance matrix $Z$ between the original features $X$ and a set of new features $\widetilde{X}$ in which each feature is modified by the mini-batches that appear before it during the learning process in an averaged way. Using this representation, we establish that the dynamics of mini-batch and full-batch gradient descent agree up to leading order with respect to the step size using the linear scaling rule. However, mini-batch gradient descent with random reshuffling exhibits a subtle dependence on the step size that a gradient flow analysis cannot detect, such as converging to a limit that depends on the step size. By comparing $Z$, a non-commutative polynomial of random matrices, with the sample covariance matrix of $X$ asymptotically, we demonstrate that batching affects the dynamics by resulting in a form of shrinkage on the spectrum.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03696
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Error dynamics of mini-batch gradient descent with random reshuffling for least squares regression
Lok, Jackie
Sonthalia, Rishi
Rebrova, Elizaveta
Machine Learning
Optimization and Control
We study the discrete dynamics of mini-batch gradient descent with random reshuffling for least squares regression. We show that the training and generalization errors depend on a sample cross-covariance matrix $Z$ between the original features $X$ and a set of new features $\widetilde{X}$ in which each feature is modified by the mini-batches that appear before it during the learning process in an averaged way. Using this representation, we establish that the dynamics of mini-batch and full-batch gradient descent agree up to leading order with respect to the step size using the linear scaling rule. However, mini-batch gradient descent with random reshuffling exhibits a subtle dependence on the step size that a gradient flow analysis cannot detect, such as converging to a limit that depends on the step size. By comparing $Z$, a non-commutative polynomial of random matrices, with the sample covariance matrix of $X$ asymptotically, we demonstrate that batching affects the dynamics by resulting in a form of shrinkage on the spectrum.
title Error dynamics of mini-batch gradient descent with random reshuffling for least squares regression
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2406.03696