Linear Gradient Prediction with Control Variates
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909892588077056 |
|---|---|
| author | Ciosek, Kamil Felicioni, Nicolò Litwin, Juan Elenter |
| author_facet | Ciosek, Kamil Felicioni, Nicolò Litwin, Juan Elenter |
| contents | We propose a new way of training neural networks, with the goal of reducing training cost. Our method uses approximate predicted gradients instead of the full gradients that require an expensive backward pass. We derive a control-variate-based technique that ensures our updates are unbiased estimates of the true gradient. Moreover, we propose a novel way to derive a predictor for the gradient inspired by the theory of the Neural Tangent Kernel. We empirically show the efficacy of the technique on a vision transformer classification task. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_05187 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Linear Gradient Prediction with Control Variates Ciosek, Kamil Felicioni, Nicolò Litwin, Juan Elenter Machine Learning We propose a new way of training neural networks, with the goal of reducing training cost. Our method uses approximate predicted gradients instead of the full gradients that require an expensive backward pass. We derive a control-variate-based technique that ensures our updates are unbiased estimates of the true gradient. Moreover, we propose a novel way to derive a predictor for the gradient inspired by the theory of the Neural Tangent Kernel. We empirically show the efficacy of the technique on a vision transformer classification task. |
| title | Linear Gradient Prediction with Control Variates |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2511.05187 |