SPLA: Block Sparse Plus Linear Attention for Long Context Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Bailin, Friedman, Dan, Lei, Tao, Wang, Chong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912861669818368
author Wang, Bailin
Friedman, Dan
Lei, Tao
Wang, Chong
author_facet Wang, Bailin
Friedman, Dan
Lei, Tao
Wang, Chong
contents Block-wise sparse attention offers significant efficiency gains for long-context modeling, yet existing methods often suffer from low selection fidelity and cumulative contextual loss by completely discarding unselected blocks. To address these limitations, we introduce Sparse Plus Linear Attention (SPLA), a framework that utilizes a selection metric derived from second-order Taylor expansions to accurately identify relevant blocks for exact attention. Instead of discarding the remaining "long tail," SPLA compresses unselected blocks into a compact recurrent state via a residual linear attention (RLA) module. Crucially, to avoid IO overhead, we derive an optimized subtraction-based formulation for RLA -- calculating the residual as the difference between global and selected linear attention -- ensuring that unselected blocks are never explicitly accessed during inference. Our experiments demonstrate that SPLA closes the performance gap in continual pretraining, surpassing dense attention models on long-context benchmarks like RULER while maintaining competitive general knowledge and reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22379
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
Wang, Bailin
Friedman, Dan
Lei, Tao
Wang, Chong
Computation and Language
Block-wise sparse attention offers significant efficiency gains for long-context modeling, yet existing methods often suffer from low selection fidelity and cumulative contextual loss by completely discarding unselected blocks. To address these limitations, we introduce Sparse Plus Linear Attention (SPLA), a framework that utilizes a selection metric derived from second-order Taylor expansions to accurately identify relevant blocks for exact attention. Instead of discarding the remaining "long tail," SPLA compresses unselected blocks into a compact recurrent state via a residual linear attention (RLA) module. Crucially, to avoid IO overhead, we derive an optimized subtraction-based formulation for RLA -- calculating the residual as the difference between global and selected linear attention -- ensuring that unselected blocks are never explicitly accessed during inference. Our experiments demonstrate that SPLA closes the performance gap in continual pretraining, surpassing dense attention models on long-context benchmarks like RULER while maintaining competitive general knowledge and reasoning capabilities.
title SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
topic Computation and Language
url https://arxiv.org/abs/2601.22379