Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908760624070656 |
|---|---|
| author | Flügel, Katharina Coquelin, Daniel Weiel, Marie Debus, Charlotte Streit, Achim Götz, Markus |
| author_facet | Flügel, Katharina Coquelin, Daniel Weiel, Marie Debus, Charlotte Streit, Achim Götz, Markus |
| contents | The gradients used to train neural networks are typically computed using backpropagation. While an efficient way to obtain exact gradients, backpropagation is computationally expensive, hinders parallelization, and is biologically implausible. Forward gradients are an approach to approximate the gradients from directional derivatives along random tangents computed by forward-mode automatic differentiation. So far, research has focused on using a single tangent per step. This paper provides an in-depth analysis of multi-tangent forward gradients and introduces an improved approach to combining the forward gradients from multiple tangents based on orthogonal projections. We demonstrate that increasing the number of tangents improves both approximation quality and optimization performance across various tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_17764 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients Flügel, Katharina Coquelin, Daniel Weiel, Marie Debus, Charlotte Streit, Achim Götz, Markus Machine Learning Artificial Intelligence The gradients used to train neural networks are typically computed using backpropagation. While an efficient way to obtain exact gradients, backpropagation is computationally expensive, hinders parallelization, and is biologically implausible. Forward gradients are an approach to approximate the gradients from directional derivatives along random tangents computed by forward-mode automatic differentiation. So far, research has focused on using a single tangent per step. This paper provides an in-depth analysis of multi-tangent forward gradients and introduces an improved approach to combining the forward gradients from multiple tangents based on orthogonal projections. We demonstrate that increasing the number of tangents improves both approximation quality and optimization performance across various tasks. |
| title | Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2410.17764 |