STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Meng, Weikang, Huo, Liangyu, Luo, Yadan, Guan, Jiawen, Zhang, Jingyi, Li, Yingjian, Zhang, Zheng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910008847892480
author Meng, Weikang
Huo, Liangyu
Luo, Yadan
Guan, Jiawen
Zhang, Jingyi
Li, Yingjian
Zhang, Zheng
author_facet Meng, Weikang
Huo, Liangyu
Luo, Yadan
Guan, Jiawen
Zhang, Jingyi
Li, Yingjian
Zhang, Zheng
contents Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local-global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training-inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02180
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
Meng, Weikang
Huo, Liangyu
Luo, Yadan
Guan, Jiawen
Zhang, Jingyi
Li, Yingjian
Zhang, Zheng
Machine Learning
Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local-global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training-inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks.
title STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
topic Machine Learning
url https://arxiv.org/abs/2602.02180