S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Qingsen, Wang, Dianyun, Wang, Yaoye, Ning, Lechen, Zhu, Sujie, Zhang, Xiaohang, Lyu, Jiaming, Ren, Linhao, Xu, Zhenbo, He, Zhaofeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912855498948608
author Ma, Qingsen
Wang, Dianyun
Wang, Yaoye
Ning, Lechen
Zhu, Sujie
Zhang, Xiaohang
Lyu, Jiaming
Ren, Linhao
Xu, Zhenbo
He, Zhaofeng
author_facet Ma, Qingsen
Wang, Dianyun
Wang, Yaoye
Ning, Lechen
Zhu, Sujie
Zhang, Xiaohang
Lyu, Jiaming
Ren, Linhao
Xu, Zhenbo
He, Zhaofeng
contents Large language models are increasingly applied to multi-document and long-form inputs, yet long-context inference remains memory- and noise-inefficient. Key-value (KV) caching scales linearly with context length, while external retrieval methods often return lexically similar but causally irrelevant passages. We present S3-Attention, a memory-first inference-time framework that treats long-context processing as attention-aligned endogenous retrieval. S3-Attention decodes transient key and query projections into top-k sparse feature identifiers using lightweight sparse autoencoders, and constructs a CPU-based inverted index mapping features to token positions or spans during a single streaming scan. This design allows the KV cache to be discarded entirely and bounds GPU memory usage by the scan chunk size. At generation time, feature co-activation is used to retrieve compact evidence spans, optionally fused with BM25 for exact lexical matching. Under a unified LongBench evaluation protocol with fixed prompting, decoding, and matched token budgets, S3-Hybrid closely matches full-context inference across multiple model families and improves robustness in several information-dense settings. We also report an engineering limitation of the current prototype, which incurs higher wall-clock latency than optimized full-KV baselines, motivating future kernel-level optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17702
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
Ma, Qingsen
Wang, Dianyun
Wang, Yaoye
Ning, Lechen
Zhu, Sujie
Zhang, Xiaohang
Lyu, Jiaming
Ren, Linhao
Xu, Zhenbo
He, Zhaofeng
Computation and Language
Machine Learning
Large language models are increasingly applied to multi-document and long-form inputs, yet long-context inference remains memory- and noise-inefficient. Key-value (KV) caching scales linearly with context length, while external retrieval methods often return lexically similar but causally irrelevant passages. We present S3-Attention, a memory-first inference-time framework that treats long-context processing as attention-aligned endogenous retrieval. S3-Attention decodes transient key and query projections into top-k sparse feature identifiers using lightweight sparse autoencoders, and constructs a CPU-based inverted index mapping features to token positions or spans during a single streaming scan. This design allows the KV cache to be discarded entirely and bounds GPU memory usage by the scan chunk size. At generation time, feature co-activation is used to retrieve compact evidence spans, optionally fused with BM25 for exact lexical matching. Under a unified LongBench evaluation protocol with fixed prompting, decoding, and matched token budgets, S3-Hybrid closely matches full-context inference across multiple model families and improves robustness in several information-dense settings. We also report an engineering limitation of the current prototype, which incurs higher wall-clock latency than optimized full-KV baselines, motivating future kernel-level optimization.
title S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.17702