FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lai, Xunhao, Lu, Jianqiao, Luo, Yao, Ma, Yiyuan, Zhou, Xun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
von: Peng, Dan, et al.
Veröffentlicht: (2025)
von: Peng, Dan, et al.
Veröffentlicht: (2025)
Enhancing Linear Attention with Residual Learning
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
von: Gu, Zhuohan, et al.
Veröffentlicht: (2024)
von: Gu, Zhuohan, et al.
Veröffentlicht: (2024)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
von: Guanzhong, Chen
Veröffentlicht: (2026)
von: Guanzhong, Chen
Veröffentlicht: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
Multipole Attention for Efficient Long Context Reasoning
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
Model Merging in Pre-training of Large Language Models
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2024)
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2024)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
von: Filipek, Adam
Veröffentlicht: (2025)
von: Filipek, Adam
Veröffentlicht: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024)
von: Lou, Chao, et al.
Veröffentlicht: (2024)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
von: Zhou, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhou, Ruijie, et al.
Veröffentlicht: (2026)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
von: Zhou, Lang, et al.
Veröffentlicht: (2026)
von: Zhou, Lang, et al.
Veröffentlicht: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
von: Ma, Xuezhe, et al.
Veröffentlicht: (2024)
von: Ma, Xuezhe, et al.
Veröffentlicht: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026) -
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
von: Peng, Dan, et al.
Veröffentlicht: (2025) -
Enhancing Linear Attention with Residual Learning
von: Lai, Xunhao, et al.
Veröffentlicht: (2025) -
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)