Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Kan, Tang, Tian, Xu, Qinyu, Gu, Yile, Zeng, Zhichen, Kadekodi, Rohan, Zhao, Liangyu, Li, Ang, Krishnamurthy, Arvind, Kasikci, Baris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
VoxServe: Streaming-Centric Serving System for Speech Language Models
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
by: Kadekodi, Rohan, et al.
Published: (2025)
by: Kadekodi, Rohan, et al.
Published: (2025)
Jenga: Responsive Tiered Memory Management without Thrashing
by: Kadekodi, Rohan, et al.
Published: (2025)
by: Kadekodi, Rohan, et al.
Published: (2025)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
by: Zhu, Kan, et al.
Published: (2024)
by: Zhu, Kan, et al.
Published: (2024)
Scalable and Accurate Application-Level Crash-Consistency Testing via Representative Testing
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
by: Zhao, Yilong, et al.
Published: (2023)
by: Zhao, Yilong, et al.
Published: (2023)
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
by: Gao, Yizhao, et al.
Published: (2024)
by: Gao, Yizhao, et al.
Published: (2024)
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025)
by: Vasylenko, Pavlo, et al.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
by: Ni, Wentao, et al.
Published: (2026)
by: Ni, Wentao, et al.
Published: (2026)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
by: Zhao, Yilong, et al.
Published: (2026)
by: Zhao, Yilong, et al.
Published: (2026)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
by: Pan, Yi, et al.
Published: (2026)
by: Pan, Yi, et al.
Published: (2026)
Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
by: Kamahori, Keisuke, et al.
Published: (2025)
by: Kamahori, Keisuke, et al.
Published: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism
by: Bu, Tao, et al.
Published: (2025)
by: Bu, Tao, et al.
Published: (2025)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
by: Krishnamurthy, Vikram
Published: (2026)
by: Krishnamurthy, Vikram
Published: (2026)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
by: Basu, Prithwish, et al.
Published: (2023)
by: Basu, Prithwish, et al.
Published: (2023)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
by: Jiang, Huiqiang, et al.
Published: (2024)
by: Jiang, Huiqiang, et al.
Published: (2024)
Mixture of In-Context Experts Enhance LLMs' Long Context Awareness
by: Lin, Hongzhan, et al.
Published: (2024)
by: Lin, Hongzhan, et al.
Published: (2024)
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
by: Tan, Haoyue, et al.
Published: (2026)
by: Tan, Haoyue, et al.
Published: (2026)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
by: Liu, Anmin, et al.
Published: (2026)
by: Liu, Anmin, et al.
Published: (2026)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs
by: Tian, Runchu, et al.
Published: (2024)
by: Tian, Runchu, et al.
Published: (2024)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
by: Zarkadas, Ioannis, et al.
Published: (2025)
by: Zarkadas, Ioannis, et al.
Published: (2025)
Attention in Constant Time: Vashista Sparse Attention for Long-Context Decoding with Exponential Guarantees
by: Nobaub, Vashista
Published: (2026)
by: Nobaub, Vashista
Published: (2026)
Similar Items
-
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025) -
VoxServe: Streaming-Centric Serving System for Speech Language Models
by: Kamahori, Keisuke, et al.
Published: (2026) -
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024) -
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
by: Kadekodi, Rohan, et al.
Published: (2025) -
Jenga: Responsive Tiered Memory Management without Thrashing
by: Kadekodi, Rohan, et al.
Published: (2025)