Saved in:
Bibliographic Details
Main Authors: Wang, Ziming, Wang, Xiang, Peng, Kailong, Qin, Lang, Kostelec, Juan Gabriel, Sourmpis, Christos, Laborieux, Axel, Guo, Qinghai
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.13680
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908834599010304
author Wang, Ziming
Wang, Xiang
Peng, Kailong
Qin, Lang
Kostelec, Juan Gabriel
Sourmpis, Christos
Laborieux, Axel
Guo, Qinghai
author_facet Wang, Ziming
Wang, Xiang
Peng, Kailong
Qin, Lang
Kostelec, Juan Gabriel
Sourmpis, Christos
Laborieux, Axel
Guo, Qinghai
contents Large Language Models (LLMs) encounter significant performance bottlenecks in long-sequence tasks due to the computational complexity and memory overhead inherent in the self-attention mechanism. To address these challenges, we introduce \textsc{AllMem}, a novel and efficient hybrid architecture that integrates Sliding Window Attention (SWA) with non-linear Test-Time Training (TTT) memory networks. \textsc{AllMem} enables models to effectively scale to ultra-long contexts while mitigating catastrophic forgetting. This approach not only overcomes the representation constraints typical of linear memory models but also significantly reduces the computational and memory footprint during long-sequence inference. Furthermore, we implement a Memory-Efficient Fine-Tuning strategy to replace standard attention layers in pre-trained models with memory-augmented sliding window layers. This framework facilitates the efficient transformation of any off-the-shelf pre-trained LLM into an \textsc{AllMem}-based architecture. Empirical evaluations confirm that our 4k window model achieves near-lossless performance on 37k LongBench with a marginal 0.83 drop compared to full attention. Furthermore, on InfiniteBench at a 128k context, our 8k window variant outperforms full attention, which validates the effectiveness of our parameterized memory in mitigating noise and maintaining robust long-range modeling without the prohibitive costs of global attention.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13680
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AllMem: A Memory-centric Recipe for Efficient Long-context Modeling
Wang, Ziming
Wang, Xiang
Peng, Kailong
Qin, Lang
Kostelec, Juan Gabriel
Sourmpis, Christos
Laborieux, Axel
Guo, Qinghai
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) encounter significant performance bottlenecks in long-sequence tasks due to the computational complexity and memory overhead inherent in the self-attention mechanism. To address these challenges, we introduce \textsc{AllMem}, a novel and efficient hybrid architecture that integrates Sliding Window Attention (SWA) with non-linear Test-Time Training (TTT) memory networks. \textsc{AllMem} enables models to effectively scale to ultra-long contexts while mitigating catastrophic forgetting. This approach not only overcomes the representation constraints typical of linear memory models but also significantly reduces the computational and memory footprint during long-sequence inference. Furthermore, we implement a Memory-Efficient Fine-Tuning strategy to replace standard attention layers in pre-trained models with memory-augmented sliding window layers. This framework facilitates the efficient transformation of any off-the-shelf pre-trained LLM into an \textsc{AllMem}-based architecture. Empirical evaluations confirm that our 4k window model achieves near-lossless performance on 37k LongBench with a marginal 0.83 drop compared to full attention. Furthermore, on InfiniteBench at a 128k context, our 8k window variant outperforms full attention, which validates the effectiveness of our parameterized memory in mitigating noise and maintaining robust long-range modeling without the prohibitive costs of global attention.
title AllMem: A Memory-centric Recipe for Efficient Long-context Modeling
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.13680