Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Yuxuan, Tan, Jianchao, Zhang, Jiaqi, Zan, Wen, Sun, Pingwei, Lu, Yifan, Sun, Yerui, Xie, Yuchen, Cai, Xunliang, Zhang, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
von: Xu, Hongtao, et al.
Veröffentlicht: (2026)
von: Xu, Hongtao, et al.
Veröffentlicht: (2026)
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
von: Sun, Pingwei, et al.
Veröffentlicht: (2026)
von: Sun, Pingwei, et al.
Veröffentlicht: (2026)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
von: Li, Jiacheng, et al.
Veröffentlicht: (2026)
von: Li, Jiacheng, et al.
Veröffentlicht: (2026)
Accelerate Speculative Decoding with Sparse Computation in Verification
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
von: Xiu, Di, et al.
Veröffentlicht: (2025)
von: Xiu, Di, et al.
Veröffentlicht: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
von: Huo, Feiye, et al.
Veröffentlicht: (2025)
von: Huo, Feiye, et al.
Veröffentlicht: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
NOSA: Native and Offloadable Sparse Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Efficient Context Scaling with LongCat ZigZag Attention
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Fine-tuning vs Prompting, Can Language Models Understand Human Values?
von: Sun, Pingwei
Veröffentlicht: (2024)
von: Sun, Pingwei
Veröffentlicht: (2024)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
von: Qin, Jiayu, et al.
Veröffentlicht: (2025)
von: Qin, Jiayu, et al.
Veröffentlicht: (2025)
A Global-Local Attention Mechanism for Relation Classification
von: Sun, Yiping
Veröffentlicht: (2024)
von: Sun, Yiping
Veröffentlicht: (2024)
Native Hybrid Attention for Efficient Sequence Modeling
von: Du, Jusen, et al.
Veröffentlicht: (2025)
von: Du, Jusen, et al.
Veröffentlicht: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
von: Song, Enxin, et al.
Veröffentlicht: (2025)
von: Song, Enxin, et al.
Veröffentlicht: (2025)
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
von: Li, Qingyuan, et al.
Veröffentlicht: (2024)
von: Li, Qingyuan, et al.
Veröffentlicht: (2024)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
von: Liu, Anmin, et al.
Veröffentlicht: (2026)
von: Liu, Anmin, et al.
Veröffentlicht: (2026)
Local-Global Attention: An Adaptive Mechanism for Multi-Scale Feature Integration
von: Shao, Yifan
Veröffentlicht: (2024)
von: Shao, Yifan
Veröffentlicht: (2024)
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
von: Li, Qingyuan, et al.
Veröffentlicht: (2024)
von: Li, Qingyuan, et al.
Veröffentlicht: (2024)
Scaling Embeddings Outperforms Scaling Experts in Language Models
von: Liu, Hong, et al.
Veröffentlicht: (2026)
von: Liu, Hong, et al.
Veröffentlicht: (2026)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
von: Shen, Zhenyi, et al.
Veröffentlicht: (2025)
von: Shen, Zhenyi, et al.
Veröffentlicht: (2025)
SALS: Sparse Attention in Latent Space for KV cache Compression
von: Mu, Junlin, et al.
Veröffentlicht: (2025)
von: Mu, Junlin, et al.
Veröffentlicht: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
SS4D: Native 4D Generative Model via Structured Spacetime Latents
von: Li, Zhibing, et al.
Veröffentlicht: (2025)
von: Li, Zhibing, et al.
Veröffentlicht: (2025)
Hierarchical Attention Graph for Scientific Document Summarization in Global and Local Level
von: Zhao, Chenlong, et al.
Veröffentlicht: (2024)
von: Zhao, Chenlong, et al.
Veröffentlicht: (2024)
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
von: Lapautre, Nicolas, et al.
Veröffentlicht: (2025)
von: Lapautre, Nicolas, et al.
Veröffentlicht: (2025)
Rectified Sparse Attention
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
Native 3D Editing with Full Attention
von: Cai, Weiwei, et al.
Veröffentlicht: (2025)
von: Cai, Weiwei, et al.
Veröffentlicht: (2025)
Unsupervised Deformable Image Registration with Local-Global Attention and Image Decomposition
von: Huang, Zhengyong, et al.
Veröffentlicht: (2026)
von: Huang, Zhengyong, et al.
Veröffentlicht: (2026)
Hilbert-Guided Sparse Local Attention
von: Li, Yunge, et al.
Veröffentlicht: (2025)
von: Li, Yunge, et al.
Veröffentlicht: (2025)
MultiMax: Sparse and Multi-Modal Attention Learning
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
TabNSA: Native Sparse Attention for Efficient Tabular Data Learning
von: Eslamian, Ali, et al.
Veröffentlicht: (2025)
von: Eslamian, Ali, et al.
Veröffentlicht: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
Stem: Rethinking Causal Information Flow in Sparse Attention
von: Niu, Lin, et al.
Veröffentlicht: (2026)
von: Niu, Lin, et al.
Veröffentlicht: (2026)
Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
von: Wang, Qiang
Veröffentlicht: (2026)
von: Wang, Qiang
Veröffentlicht: (2026)
A Global-Local Graph Attention Network for Traffic Forecasting
von: Zhang, Tianchi
Veröffentlicht: (2026)
von: Zhang, Tianchi
Veröffentlicht: (2026)
Ähnliche Einträge
-
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026) -
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
von: Xu, Hongtao, et al.
Veröffentlicht: (2026) -
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
von: Sun, Pingwei, et al.
Veröffentlicht: (2026) -
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
von: Li, Jiacheng, et al.
Veröffentlicht: (2026) -
Accelerate Speculative Decoding with Sparse Computation in Verification
von: Wang, Jikai, et al.
Veröffentlicht: (2025)