| _version_ | 1866901414306906112 |
|---|---|
| author | Mahendrakar, Pranay |
| author_facet | Mahendrakar, Pranay |
| contents | <p>State-space models did not merely "open a door" beyond attention; they walked through it. Mamba, Mamba-2, Samba,<br>Jamba, Granite 4, Hymba, and a growing family of hybrid attention-SSM architectures are now in production language models,<br>demonstrating that linear-time alternatives to attention can match transformers on language modelling at competitive scales.<br>The simultaneous expansion of pure-attention context windows — Gemini 1.5 and successors handling up to 10 million tokens<br>with near-perfect needle-in-haystack recall — has changed the empirical landscape that motivated SSM research in the first<br>place. The popular framing of "memory architectures beyond attention" has not kept up with this. This paper makes three<br>claims. First, the word "memory" in the long-context discussion conflates four distinct concepts — architectural state, context<br>window, external retrieval, and persistent agent memory — each with different scaling properties and different research<br>questions. Second, the empirical picture is more nuanced than either the SSM-replaces-attention or the attention-is-enough<br>framings: SSMs win on very long passive recall and inference efficiency, transformers win on complex reasoning, hybrids win<br>in deployment, and the choice between them is task-dependent. Third, the genuine open frontier is reasoning at long context<br>— not retrieval, which is largely solved — and benchmarks like MathHay (51% accuracy at 128K tokens for Gemini-1.5-Pro)<br>make the gap quantitatively visible. We propose a research agenda focused on architecture-task fit, hierarchical multi-scale<br>memory, and reasoning-aware long-context evaluation.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19855022 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Memory Architectures Beyond Attention: Disambiguating Four Memory Concepts and the Long-Context Reasoning Frontier Mahendrakar, Pranay <p>State-space models did not merely "open a door" beyond attention; they walked through it. Mamba, Mamba-2, Samba,<br>Jamba, Granite 4, Hymba, and a growing family of hybrid attention-SSM architectures are now in production language models,<br>demonstrating that linear-time alternatives to attention can match transformers on language modelling at competitive scales.<br>The simultaneous expansion of pure-attention context windows — Gemini 1.5 and successors handling up to 10 million tokens<br>with near-perfect needle-in-haystack recall — has changed the empirical landscape that motivated SSM research in the first<br>place. The popular framing of "memory architectures beyond attention" has not kept up with this. This paper makes three<br>claims. First, the word "memory" in the long-context discussion conflates four distinct concepts — architectural state, context<br>window, external retrieval, and persistent agent memory — each with different scaling properties and different research<br>questions. Second, the empirical picture is more nuanced than either the SSM-replaces-attention or the attention-is-enough<br>framings: SSMs win on very long passive recall and inference efficiency, transformers win on complex reasoning, hybrids win<br>in deployment, and the choice between them is task-dependent. Third, the genuine open frontier is reasoning at long context<br>— not retrieval, which is largely solved — and benchmarks like MathHay (51% accuracy at 128K tokens for Gemini-1.5-Pro)<br>make the gap quantitatively visible. We propose a research agenda focused on architecture-task fit, hierarchical multi-scale<br>memory, and reasoning-aware long-context evaluation.</p> |
| title | Memory Architectures Beyond Attention: Disambiguating Four Memory Concepts and the Long-Context Reasoning Frontier |
| url | https://doi.org/10.5281/zenodo.19855022 |