FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Ran, Jiang, Youhe, Chen, Zhuoming, Mai, Haohui, Chen, Beidi, Yuan, Binhang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2023)
by: Jiang, Youhe, et al.
Published: (2023)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
by: Liang, Yan, et al.
Published: (2026)
by: Liang, Yan, et al.
Published: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024)
by: Yan, Ran, et al.
Published: (2024)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
by: Wang, Huanyu, et al.
Published: (2026)
by: Wang, Huanyu, et al.
Published: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
Cascadia: An Efficient Cascade Serving System for Large Language Models
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
by: Tong, Chris, et al.
Published: (2025)
by: Tong, Chris, et al.
Published: (2025)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
by: Chen, Shaoyuan, et al.
Published: (2024)
by: Chen, Shaoyuan, et al.
Published: (2024)
Reproduction Research of FSA-Benchmark
by: Ludolf, Joshua, et al.
Published: (2024)
by: Ludolf, Joshua, et al.
Published: (2024)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
by: He, Guoliang, et al.
Published: (2025)
by: He, Guoliang, et al.
Published: (2025)
Efficient Long-context Language Model Training by Core Attention Disaggregation
by: Zhuang, Yonghao, et al.
Published: (2025)
by: Zhuang, Yonghao, et al.
Published: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
Fused3S: Fast Sparse Attention on Tensor Cores
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023)
by: Zhao, Minjun, et al.
Published: (2023)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native
by: Lu, Yao, et al.
Published: (2024)
by: Lu, Yao, et al.
Published: (2024)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
Optimizing the Optimal Weighted Average: Efficient Distributed Sparse Classification
by: Lu, Fred, et al.
Published: (2024)
by: Lu, Fred, et al.
Published: (2024)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2024)
by: Chen, Shuangyi, et al.
Published: (2024)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
by: Pan, Xinglin, et al.
Published: (2024)
by: Pan, Xinglin, et al.
Published: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
by: Jiang, Yinsicheng, et al.
Published: (2025)
by: Jiang, Yinsicheng, et al.
Published: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
by: Jiang, Yinsicheng, et al.
Published: (2024)
by: Jiang, Yinsicheng, et al.
Published: (2024)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
by: Wang, Yujie, et al.
Published: (2023)
by: Wang, Yujie, et al.
Published: (2023)
Stabilizing Decentralized Federated Fine-Tuning via Topology-Aware Alternating LoRA
by: Wang, Xiaoyu, et al.
Published: (2026)
by: Wang, Xiaoyu, et al.
Published: (2026)
High-Dimensional Distributed Sparse Classification with Scalable Communication-Efficient Global Updates
by: Lu, Fred, et al.
Published: (2024)
by: Lu, Fred, et al.
Published: (2024)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
by: Chang, Li-Wen, et al.
Published: (2024)
by: Chang, Li-Wen, et al.
Published: (2024)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Tackling Resource-Constrained and Data-Heterogeneity in Federated Learning with Double-Weight Sparse Pack
by: Yang, Qiantao, et al.
Published: (2026)
by: Yang, Qiantao, et al.
Published: (2026)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
by: Liang, Yunkai, et al.
Published: (2025)
by: Liang, Yunkai, et al.
Published: (2025)
BatchWeave: A Consistent Object-Store-Native Data Plane for Large Foundation Model Training
by: Sun, Ting, et al.
Published: (2026)
by: Sun, Ting, et al.
Published: (2026)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
by: Chen, Chuyan, et al.
Published: (2025)
by: Chen, Chuyan, et al.
Published: (2025)
Similar Items
-
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2023) -
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
by: Yan, Ran, et al.
Published: (2025) -
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2025) -
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
by: Liang, Yan, et al.
Published: (2026) -
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024)