Reservoir-Augmented Attention: Combining Echo State Networks with Transformer Attention for Efficient Language Modelling

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Josh, Taylor
Natura: Recurso digital
Lingua:En
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902089286811648
author Josh, Taylor
author_facet Josh, Taylor
contents <p>We present two hybrid architectures that combine Echo State Network (ESN) reservoirs with transformer-style attention for parameter-efficient character-level language modelling. <em>Fixed-KV Attention</em> replaces learned K/V projections by fixed random linear maps of reservoir states, while <em>Node Attention</em> reframes attention as a per-step, query-gated readout over reservoir nodes, reducing attention complexity from O(T²) to O(R) per step. On TinyShakespeare, Node Attention achieves the best validation loss (1.969), beating an off-the-shelf transformer and the AERC literature approach while training at ≈21.8k tokens/s on CPU and using only 347k trained parameters. Our results show that rich reservoir dynamics plus a query-gated node readout provide a compact, efficient alternative to learned K/V projections and point to a practical route for long-context, efficient language modelling.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18903774
institution Zenodo
language enc
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Reservoir-Augmented Attention: Combining Echo State Networks with Transformer Attention for Efficient Language Modelling
Josh, Taylor
Echo State Networks
Reservoir Computing
Attention
Large language model
LLM
NLP
<p>We present two hybrid architectures that combine Echo State Network (ESN) reservoirs with transformer-style attention for parameter-efficient character-level language modelling. <em>Fixed-KV Attention</em> replaces learned K/V projections by fixed random linear maps of reservoir states, while <em>Node Attention</em> reframes attention as a per-step, query-gated readout over reservoir nodes, reducing attention complexity from O(T²) to O(R) per step. On TinyShakespeare, Node Attention achieves the best validation loss (1.969), beating an off-the-shelf transformer and the AERC literature approach while training at ≈21.8k tokens/s on CPU and using only 347k trained parameters. Our results show that rich reservoir dynamics plus a query-gated node readout provide a compact, efficient alternative to learned K/V projections and point to a practical route for long-context, efficient language modelling.</p>
title Reservoir-Augmented Attention: Combining Echo State Networks with Transformer Attention for Efficient Language Modelling
topic Echo State Networks
Reservoir Computing
Attention
Large language model
LLM
NLP
url https://doi.org/10.5281/zenodo.18903774