A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiu, Di, Tang, Hongyin, Rong, Bolin, Yan, Lizhi, Wang, Jingang, Lu, Yifan, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
NOSA: Native and Offloadable Sparse Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Node Classification via Semantic-Structural Attention-Enhanced Graph Convolutional Networks
von: Zhu, Hongyin
Veröffentlicht: (2024)
von: Zhu, Hongyin
Veröffentlicht: (2024)
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
von: Bai, Yang, et al.
Veröffentlicht: (2026)
von: Bai, Yang, et al.
Veröffentlicht: (2026)
Length Desensitization in Direct Preference Optimization
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
von: Li, Jiacheng, et al.
Veröffentlicht: (2026)
von: Li, Jiacheng, et al.
Veröffentlicht: (2026)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
von: Wang, Junxuan, et al.
Veröffentlicht: (2025)
von: Wang, Junxuan, et al.
Veröffentlicht: (2025)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
Unraveling the Mystery of Scaling Laws: Part I
von: Su, Hui, et al.
Veröffentlicht: (2024)
von: Su, Hui, et al.
Veröffentlicht: (2024)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025)
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025)
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
Construct, Align, and Reason: Large Ontology Models for Enterprise Knowledge Management
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
von: Bai, Yushi, et al.
Veröffentlicht: (2026)
von: Bai, Yushi, et al.
Veröffentlicht: (2026)
Native Hybrid Attention for Efficient Sequence Modeling
von: Du, Jusen, et al.
Veröffentlicht: (2025)
von: Du, Jusen, et al.
Veröffentlicht: (2025)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
von: Chen, Lida, et al.
Veröffentlicht: (2025)
von: Chen, Lida, et al.
Veröffentlicht: (2025)
Block Sparse Flash Attention
von: Ohayon, Daniel, et al.
Veröffentlicht: (2025)
von: Ohayon, Daniel, et al.
Veröffentlicht: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
AdaSplash: Adaptive Sparse Flash Attention
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2025)
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2025)
Scaling Linear Attention with Sparse State Expansion
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
von: Tzachristas, Georgios, et al.
Veröffentlicht: (2025)
von: Tzachristas, Georgios, et al.
Veröffentlicht: (2025)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025)
von: He, Mutian, et al.
Veröffentlicht: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
Sparse Attention across Multiple-context KV Cache
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
AdaSplash-2: Faster Differentiable Sparse Attention
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2026)
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2026)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024) -
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025) -
NOSA: Native and Offloadable Sparse Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025) -
Node Classification via Semantic-Structural Attention-Enhanced Graph Convolutional Networks
von: Zhu, Hongyin
Veröffentlicht: (2024) -
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)