Convergence guarantees for forward gradient descent in the linear regression model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bos, Thijs, Schmidt-Hieber, Johannes
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917699138879488
author Bos, Thijs
Schmidt-Hieber, Johannes
author_facet Bos, Thijs
Schmidt-Hieber, Johannes
contents Renewed interest in the relationship between artificial and biological neural networks motivates the study of gradient-free methods. Considering the linear regression model with random design, we theoretically analyze in this work the biologically motivated (weight-perturbed) forward gradient scheme that is based on random linear combination of the gradient. If d denotes the number of parameters and k the number of samples, we prove that the mean squared error of this method converges for $k\gtrsim d^2\log(d)$ with rate $d^2\log(d)/k.$ Compared to the dimension dependence d for stochastic gradient descent, an additional factor $d\log(d)$ occurs.
format Preprint
id arxiv_https___arxiv_org_abs_2309_15001
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Convergence guarantees for forward gradient descent in the linear regression model
Bos, Thijs
Schmidt-Hieber, Johannes
Statistics Theory
Neural and Evolutionary Computing
Primary: 62L20, secondary: 62J05
Renewed interest in the relationship between artificial and biological neural networks motivates the study of gradient-free methods. Considering the linear regression model with random design, we theoretically analyze in this work the biologically motivated (weight-perturbed) forward gradient scheme that is based on random linear combination of the gradient. If d denotes the number of parameters and k the number of samples, we prove that the mean squared error of this method converges for $k\gtrsim d^2\log(d)$ with rate $d^2\log(d)/k.$ Compared to the dimension dependence d for stochastic gradient descent, an additional factor $d\log(d)$ occurs.
title Convergence guarantees for forward gradient descent in the linear regression model
topic Statistics Theory
Neural and Evolutionary Computing
Primary: 62L20, secondary: 62J05
url https://arxiv.org/abs/2309.15001