Fast Training of Recurrent Neural Networks with Stationary State Feedbacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Caillon, Paul, Fagnou, Erwan, Allauzen, Alexandre
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917971832602624
author Caillon, Paul
Fagnou, Erwan
Allauzen, Alexandre
author_facet Caillon, Paul
Fagnou, Erwan
Allauzen, Alexandre
contents Recurrent neural networks (RNNs) have recently demonstrated strong performance and faster inference than Transformers at comparable parameter budgets. However, the recursive gradient computation with the backpropagation through time (or BPTT) algorithm remains the major computational bottleneck. In this work, we propose a novel method that replaces BPTT with a fixed gradient feedback mechanism, yielding an efficient approximation of the exact gradient propagation based on the assumption of time stationarity. Our approach leverages state-space model (SSM) principles to define a structured feedback matrix that directly propagates gradients from future time steps. This formulation bypasses the need for recursive gradient backpropagation, significantly reducing training overhead while preserving the network's ability to capture long-term dependencies. The experiments on language modeling benchmarks exhibit competitive perplexity scores, while significantly reducing the training costs. These promising results suggest that designing a feedback method like an SSM can fully exploit the efficiency advantages of RNNs for many practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fast Training of Recurrent Neural Networks with Stationary State Feedbacks
Caillon, Paul
Fagnou, Erwan
Allauzen, Alexandre
Machine Learning
Artificial Intelligence
Recurrent neural networks (RNNs) have recently demonstrated strong performance and faster inference than Transformers at comparable parameter budgets. However, the recursive gradient computation with the backpropagation through time (or BPTT) algorithm remains the major computational bottleneck. In this work, we propose a novel method that replaces BPTT with a fixed gradient feedback mechanism, yielding an efficient approximation of the exact gradient propagation based on the assumption of time stationarity. Our approach leverages state-space model (SSM) principles to define a structured feedback matrix that directly propagates gradients from future time steps. This formulation bypasses the need for recursive gradient backpropagation, significantly reducing training overhead while preserving the network's ability to capture long-term dependencies. The experiments on language modeling benchmarks exhibit competitive perplexity scores, while significantly reducing the training costs. These promising results suggest that designing a feedback method like an SSM can fully exploit the efficiency advantages of RNNs for many practical applications.
title Fast Training of Recurrent Neural Networks with Stationary State Feedbacks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.23104