Saved in:
Bibliographic Details
Main Authors: Kiruluta, Andrew, Raju, Preethi, Burity, Priscilla
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.01963
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913871855353856
author Kiruluta, Andrew
Raju, Preethi
Burity, Priscilla
author_facet Kiruluta, Andrew
Raju, Preethi
Burity, Priscilla
contents We present a novel non attention based architecture for large language models (LLMs) that efficiently handles very long context windows, on the order of hundreds of thousands to potentially millions of tokens. Unlike traditional Transformer designs, which suffer from quadratic memory and computation overload due to the nature of the self attention mechanism, our model avoids token to token attention entirely. Instead, it combines the following complementary components: State Space blocks (inspired by S4) that learn continuous time convolution kernels and scale near linearly with sequence length, Multi Resolution Convolution layers that capture local context at different dilation levels, a lightweight Recurrent Supervisor to maintain a global hidden state across sequential chunks, and Retrieval Augmented External Memory that stores and retrieves high-level chunk embeddings without reintroducing quadratic operations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
Kiruluta, Andrew
Raju, Preethi
Burity, Priscilla
Machine Learning
Computation and Language
We present a novel non attention based architecture for large language models (LLMs) that efficiently handles very long context windows, on the order of hundreds of thousands to potentially millions of tokens. Unlike traditional Transformer designs, which suffer from quadratic memory and computation overload due to the nature of the self attention mechanism, our model avoids token to token attention entirely. Instead, it combines the following complementary components: State Space blocks (inspired by S4) that learn continuous time convolution kernels and scale near linearly with sequence length, Multi Resolution Convolution layers that capture local context at different dilation levels, a lightweight Recurrent Supervisor to maintain a global hidden state across sequential chunks, and Retrieval Augmented External Memory that stores and retrieves high-level chunk embeddings without reintroducing quadratic operations.
title Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.01963