Convergence guarantees for forward gradient descent in the linear regression model
Fuente:
arXiv
Guardado en:
| Autores principales: | , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917699138879488 |
|---|---|
| author | Bos, Thijs Schmidt-Hieber, Johannes |
| author_facet | Bos, Thijs Schmidt-Hieber, Johannes |
| contents | Renewed interest in the relationship between artificial and biological neural networks motivates the study of gradient-free methods. Considering the linear regression model with random design, we theoretically analyze in this work the biologically motivated (weight-perturbed) forward gradient scheme that is based on random linear combination of the gradient. If d denotes the number of parameters and k the number of samples, we prove that the mean squared error of this method converges for $k\gtrsim d^2\log(d)$ with rate $d^2\log(d)/k.$ Compared to the dimension dependence d for stochastic gradient descent, an additional factor $d\log(d)$ occurs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_15001 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Convergence guarantees for forward gradient descent in the linear regression model Bos, Thijs Schmidt-Hieber, Johannes Statistics Theory Neural and Evolutionary Computing Primary: 62L20, secondary: 62J05 Renewed interest in the relationship between artificial and biological neural networks motivates the study of gradient-free methods. Considering the linear regression model with random design, we theoretically analyze in this work the biologically motivated (weight-perturbed) forward gradient scheme that is based on random linear combination of the gradient. If d denotes the number of parameters and k the number of samples, we prove that the mean squared error of this method converges for $k\gtrsim d^2\log(d)$ with rate $d^2\log(d)/k.$ Compared to the dimension dependence d for stochastic gradient descent, an additional factor $d\log(d)$ occurs. |
| title | Convergence guarantees for forward gradient descent in the linear regression model |
| topic | Statistics Theory Neural and Evolutionary Computing Primary: 62L20, secondary: 62J05 |
| url | https://arxiv.org/abs/2309.15001 |