Gradient-free training of recurrent neural networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bolager, Erik Lien, Cukarska, Ana, Burak, Iryna, Monfared, Zahra, Dietrich, Felix
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910804213760000
author Bolager, Erik Lien
Cukarska, Ana
Burak, Iryna
Monfared, Zahra
Dietrich, Felix
author_facet Bolager, Erik Lien
Cukarska, Ana
Burak, Iryna
Monfared, Zahra
Dietrich, Felix
contents Recurrent neural networks are a successful neural architecture for many time-dependent problems, including time series analysis, forecasting, and modeling of dynamical systems. Training such networks with backpropagation through time is a notoriously difficult problem because their loss gradients tend to explode or vanish. In this contribution, we introduce a computational approach to construct all weights and biases of a recurrent neural network without using gradient-based methods. The approach is based on a combination of random feature networks and Koopman operator theory for dynamical systems. The hidden parameters of a single recurrent block are sampled at random, while the outer weights are constructed using extended dynamic mode decomposition. This approach alleviates all problems with backpropagation commonly related to recurrent networks. The connection to Koopman operator theory also allows us to start using results in this area to analyze recurrent neural networks. In computational experiments on time series, forecasting for chaotic dynamical systems, and control problems, as well as on weather data, we observe that the training time and forecasting accuracy of the recurrent neural networks we construct are improved when compared to commonly used gradient-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gradient-free training of recurrent neural networks
Bolager, Erik Lien
Cukarska, Ana
Burak, Iryna
Monfared, Zahra
Dietrich, Felix
Machine Learning
Numerical Analysis
Recurrent neural networks are a successful neural architecture for many time-dependent problems, including time series analysis, forecasting, and modeling of dynamical systems. Training such networks with backpropagation through time is a notoriously difficult problem because their loss gradients tend to explode or vanish. In this contribution, we introduce a computational approach to construct all weights and biases of a recurrent neural network without using gradient-based methods. The approach is based on a combination of random feature networks and Koopman operator theory for dynamical systems. The hidden parameters of a single recurrent block are sampled at random, while the outer weights are constructed using extended dynamic mode decomposition. This approach alleviates all problems with backpropagation commonly related to recurrent networks. The connection to Koopman operator theory also allows us to start using results in this area to analyze recurrent neural networks. In computational experiments on time series, forecasting for chaotic dynamical systems, and control problems, as well as on weather data, we observe that the training time and forecasting accuracy of the recurrent neural networks we construct are improved when compared to commonly used gradient-based methods.
title Gradient-free training of recurrent neural networks
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2410.23467