Reservoir-Augmented Attention: Combining Echo State Networks with Transformer Attention for Efficient Language Modelling
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Lingua: | En |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902089286811648 |
|---|---|
| author | Josh, Taylor |
| author_facet | Josh, Taylor |
| contents | <p>We present two hybrid architectures that combine Echo State Network (ESN) reservoirs with transformer-style attention for parameter-efficient character-level language modelling. <em>Fixed-KV Attention</em> replaces learned K/V projections by fixed random linear maps of reservoir states, while <em>Node Attention</em> reframes attention as a per-step, query-gated readout over reservoir nodes, reducing attention complexity from O(T²) to O(R) per step. On TinyShakespeare, Node Attention achieves the best validation loss (1.969), beating an off-the-shelf transformer and the AERC literature approach while training at ≈21.8k tokens/s on CPU and using only 347k trained parameters. Our results show that rich reservoir dynamics plus a query-gated node readout provide a compact, efficient alternative to learned K/V projections and point to a practical route for long-context, efficient language modelling.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18903774 |
| institution | Zenodo |
| language | enc |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Reservoir-Augmented Attention: Combining Echo State Networks with Transformer Attention for Efficient Language Modelling Josh, Taylor Echo State Networks Reservoir Computing Attention Large language model LLM NLP <p>We present two hybrid architectures that combine Echo State Network (ESN) reservoirs with transformer-style attention for parameter-efficient character-level language modelling. <em>Fixed-KV Attention</em> replaces learned K/V projections by fixed random linear maps of reservoir states, while <em>Node Attention</em> reframes attention as a per-step, query-gated readout over reservoir nodes, reducing attention complexity from O(T²) to O(R) per step. On TinyShakespeare, Node Attention achieves the best validation loss (1.969), beating an off-the-shelf transformer and the AERC literature approach while training at ≈21.8k tokens/s on CPU and using only 347k trained parameters. Our results show that rich reservoir dynamics plus a query-gated node readout provide a compact, efficient alternative to learned K/V projections and point to a practical route for long-context, efficient language modelling.</p> |
| title | Reservoir-Augmented Attention: Combining Echo State Networks with Transformer Attention for Efficient Language Modelling |
| topic | Echo State Networks Reservoir Computing Attention Large language model LLM NLP |
| url | https://doi.org/10.5281/zenodo.18903774 |