ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Qiuyang, Zhou, Kai, Tang, Ding, Lu, Kai, Li, Cheng, Yang, Zhenyu, Xu, Peng, Wan, Jiguang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
di: Ma, Da, et al.
Pubblicazione: (2024)
di: Ma, Da, et al.
Pubblicazione: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
di: Tao, Wei, et al.
Pubblicazione: (2026)
di: Tao, Wei, et al.
Pubblicazione: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2026)
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2026)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2024)
di: Feng, Yuan, et al.
Pubblicazione: (2024)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
Efficient Long-Context LLM Inference via KV Cache Clustering
di: Hu, Jie, et al.
Pubblicazione: (2025)
di: Hu, Jie, et al.
Pubblicazione: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
di: Zhang, Huawei, et al.
Pubblicazione: (2025)
di: Zhang, Huawei, et al.
Pubblicazione: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
di: Zuo, Youhui, et al.
Pubblicazione: (2025)
di: Zuo, Youhui, et al.
Pubblicazione: (2025)
KV Cache Offloading for Context-Intensive Tasks
di: Bocharnikov, Andrey, et al.
Pubblicazione: (2026)
di: Bocharnikov, Andrey, et al.
Pubblicazione: (2026)
Hammer: Towards Efficient Hot-Cold Data Identification via Online Learning
di: Lu, Kai, et al.
Pubblicazione: (2024)
di: Lu, Kai, et al.
Pubblicazione: (2024)
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
di: Xu, Yichun, et al.
Pubblicazione: (2026)
di: Xu, Yichun, et al.
Pubblicazione: (2026)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
di: Wu, Haoyi, et al.
Pubblicazione: (2024)
di: Wu, Haoyi, et al.
Pubblicazione: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
di: Li, Hanchen, et al.
Pubblicazione: (2025)
di: Li, Hanchen, et al.
Pubblicazione: (2025)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
Taming the Fragility of KV Cache Eviction in LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025)
di: Yu, Bohan, et al.
Pubblicazione: (2025)
Online Scheduling for LLM Inference with KV Cache Constraints
di: Jaillet, Patrick, et al.
Pubblicazione: (2025)
di: Jaillet, Patrick, et al.
Pubblicazione: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
di: Zeng, Wenxuan, et al.
Pubblicazione: (2025)
di: Zeng, Wenxuan, et al.
Pubblicazione: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
di: Yang, Yifei, et al.
Pubblicazione: (2024)
di: Yang, Yifei, et al.
Pubblicazione: (2024)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
di: Huang, Kai, et al.
Pubblicazione: (2025)
di: Huang, Kai, et al.
Pubblicazione: (2025)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
di: Wu, Guanlong, et al.
Pubblicazione: (2026)
di: Wu, Guanlong, et al.
Pubblicazione: (2026)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
di: Wan, Zhongwei, et al.
Pubblicazione: (2025)
di: Wan, Zhongwei, et al.
Pubblicazione: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
di: Chang, Chi-Chih, et al.
Pubblicazione: (2025)
di: Chang, Chi-Chih, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
di: Ma, Da, et al.
Pubblicazione: (2024) -
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
di: Tao, Wei, et al.
Pubblicazione: (2026) -
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025) -
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
di: Liu, Yuhan, et al.
Pubblicazione: (2025) -
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
di: Sharma, Akshat, et al.
Pubblicazione: (2024)