AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Di, Wang, Ruitian, Chen, Chen, Gong, Mingliang, Yuan, Yongjie, Zhao, Han, Feng, Yu, Chen, Quan, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage
von: Lin, Junqing, et al.
Veröffentlicht: (2025)
von: Lin, Junqing, et al.
Veröffentlicht: (2025)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
von: Kong, Jie, et al.
Veröffentlicht: (2025)
von: Kong, Jie, et al.
Veröffentlicht: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
von: Ni, Xiaobing, et al.
Veröffentlicht: (2024)
von: Ni, Xiaobing, et al.
Veröffentlicht: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
von: Hu, Zhen, et al.
Veröffentlicht: (2025)
von: Hu, Zhen, et al.
Veröffentlicht: (2025)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
von: Xu, Ao, et al.
Veröffentlicht: (2025)
von: Xu, Ao, et al.
Veröffentlicht: (2025)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
Stochastic Sparse Attention for Memory-Bound Inference
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
Efficient Long Context Fine-tuning with Chunk Flow
von: Yuan, Xiulong, et al.
Veröffentlicht: (2025)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2025)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
Justitia: Fair and Efficient Scheduling of Task-parallel LLM Agents with Selective Pampering
von: Yang, Mingyan, et al.
Veröffentlicht: (2025)
von: Yang, Mingyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026) -
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025) -
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025) -
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024) -
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)