Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yuxuan, Tan, Jianchao, Zhang, Jiaqi, Zan, Wen, Sun, Pingwei, Lu, Yifan, Sun, Yerui, Xie, Yuchen, Cai, Xunliang, Zhang, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911246445445120
author Hu, Yuxuan
Tan, Jianchao
Zhang, Jiaqi
Zan, Wen
Sun, Pingwei
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
Zhang, Jing
author_facet Hu, Yuxuan
Tan, Jianchao
Zhang, Jiaqi
Zan, Wen
Sun, Pingwei
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
Zhang, Jing
contents In this work, we conduct a systematic analysis of Native Sparse Attention (NSA) and propose targeted improvements that enhance long-context modeling. A key insight is that alternating between local (sliding-window) and global (compression, selective) attention across layers, rather than using fixed patterns, enables more effective propagation of long-range dependencies and substantially boosts performance on long-sequence tasks. Meanwhile, we further refine NSA's branches with Latent Attention that the sliding-window branch is enhanced with Multi-head Latent Attention (MLA) while compression and selective branches adopt Group-head Latent Attention (GLA). These changes reduce KV-cache memory by 50\% versus NSA while improving the model's common-sense reasoning and long-text understanding capabilities. Experiments on models from 340M to 1.3B parameters (trained on 15B and 100B tokens) show our method matches or exceeds full attention and native sparse attention in both common-sense reasoning and long-context understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
Hu, Yuxuan
Tan, Jianchao
Zhang, Jiaqi
Zan, Wen
Sun, Pingwei
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
Zhang, Jing
Computation and Language
In this work, we conduct a systematic analysis of Native Sparse Attention (NSA) and propose targeted improvements that enhance long-context modeling. A key insight is that alternating between local (sliding-window) and global (compression, selective) attention across layers, rather than using fixed patterns, enables more effective propagation of long-range dependencies and substantially boosts performance on long-sequence tasks. Meanwhile, we further refine NSA's branches with Latent Attention that the sliding-window branch is enhanced with Multi-head Latent Attention (MLA) while compression and selective branches adopt Group-head Latent Attention (GLA). These changes reduce KV-cache memory by 50\% versus NSA while improving the model's common-sense reasoning and long-text understanding capabilities. Experiments on models from 340M to 1.3B parameters (trained on 15B and 100B tokens) show our method matches or exceeds full attention and native sparse attention in both common-sense reasoning and long-context understanding tasks.
title Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
topic Computation and Language
url https://arxiv.org/abs/2511.00819