SparseAccelerate: Efficient Long-Context Inference for Mid-Range GPUs
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Vo, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
von: Synk, Ryan, et al.
Veröffentlicht: (2025)
von: Synk, Ryan, et al.
Veröffentlicht: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
Squeezed Attention: Accelerating Long Context Length LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
von: Ge, Suyu, et al.
Veröffentlicht: (2024)
von: Ge, Suyu, et al.
Veröffentlicht: (2024)
Transformer Layer Injection: A Novel Approach for Efficient Upscaling of Large Language Models
von: Vo, James
Veröffentlicht: (2024)
von: Vo, James
Veröffentlicht: (2024)
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
von: Long, Lingkun, et al.
Veröffentlicht: (2025)
von: Long, Lingkun, et al.
Veröffentlicht: (2025)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
von: Long, Lingkun, et al.
Veröffentlicht: (2026)
von: Long, Lingkun, et al.
Veröffentlicht: (2026)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
von: Peng, Dan, et al.
Veröffentlicht: (2025)
von: Peng, Dan, et al.
Veröffentlicht: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
von: Shao, Wei, et al.
Veröffentlicht: (2025)
von: Shao, Wei, et al.
Veröffentlicht: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024)
von: Lou, Chao, et al.
Veröffentlicht: (2024)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
von: Ai, Xuan, et al.
Veröffentlicht: (2026)
von: Ai, Xuan, et al.
Veröffentlicht: (2026)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
Efficient Second-Order Neural Network Optimization via Adaptive Trust Region Methods
von: Vo, James
Veröffentlicht: (2024)
von: Vo, James
Veröffentlicht: (2024)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
von: Liu, Siran, et al.
Veröffentlicht: (2026)
von: Liu, Siran, et al.
Veröffentlicht: (2026)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
von: Zhang, Hanzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hanzhi, et al.
Veröffentlicht: (2025)
Long-Context Generalization with Sparse Attention
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
Lag-Relative Sparse Attention In Long Context Training
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2024)
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
von: Son, Youngjun, et al.
Veröffentlicht: (2025)
von: Son, Youngjun, et al.
Veröffentlicht: (2025)
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
von: Deng, Haoran, et al.
Veröffentlicht: (2025)
von: Deng, Haoran, et al.
Veröffentlicht: (2025)
Inference Scaling for Long-Context Retrieval Augmented Generation
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025) -
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
von: Synk, Ryan, et al.
Veröffentlicht: (2025) -
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
von: Yan, Siyuan, et al.
Veröffentlicht: (2025) -
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)