KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Minsik, Rastegari, Mohammad, Naik, Devang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
by: Li, Weiqing, et al.
Published: (2025)
by: Li, Weiqing, et al.
Published: (2025)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
by: Zhao, Bingzhe, et al.
Published: (2025)
by: Zhao, Bingzhe, et al.
Published: (2025)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
by: Zhang, Huawei, et al.
Published: (2025)
by: Zhang, Huawei, et al.
Published: (2025)
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
Leyline: KV Cache Directives for Agentic Inference
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Regulating Branch Parallelism in LLM Serving
by: Gandhi, Swapnil, et al.
Published: (2026)
by: Gandhi, Swapnil, et al.
Published: (2026)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
by: Kim, Han-Byul, et al.
Published: (2025)
by: Kim, Han-Byul, et al.
Published: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
by: Nian, Sean, et al.
Published: (2026)
by: Nian, Sean, et al.
Published: (2026)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
by: Tu, Dezhan, et al.
Published: (2024)
by: Tu, Dezhan, et al.
Published: (2024)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
by: Patel, Ishan, et al.
Published: (2026)
by: Patel, Ishan, et al.
Published: (2026)
A Survey on Large Language Model Acceleration based on KV Cache Management
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
by: Li, Yufei, et al.
Published: (2025)
by: Li, Yufei, et al.
Published: (2025)
PiKV: KV Cache Management System for Mixture of Experts
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
by: Liu, Man, et al.
Published: (2026)
by: Liu, Man, et al.
Published: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
by: Xie, Jincheng, et al.
Published: (2026)
by: Xie, Jincheng, et al.
Published: (2026)
Accelerating LLM Inference with Precomputed Query Storage
by: Park, Jay H., et al.
Published: (2025)
by: Park, Jay H., et al.
Published: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
by: Li, Wanqian, et al.
Published: (2026)
by: Li, Wanqian, et al.
Published: (2026)
Context Parallelism for Scalable Million-Token Inference
by: Yang, Amy, et al.
Published: (2024)
by: Yang, Amy, et al.
Published: (2024)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
by: Xiang, Xingyu, et al.
Published: (2025)
by: Xiang, Xingyu, et al.
Published: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
MatKV: Trading Compute for Flash Storage in LLM Inference
by: Shin, Kun-Woo, et al.
Published: (2025)
by: Shin, Kun-Woo, et al.
Published: (2025)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
by: Timor, Nadav, et al.
Published: (2024)
by: Timor, Nadav, et al.
Published: (2024)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
by: Wang, Tuowei, et al.
Published: (2025)
by: Wang, Tuowei, et al.
Published: (2025)
From Prompts to Performance: Evaluating LLMs for Task-based Parallel Code Generation
by: Bantel, Linus, et al.
Published: (2026)
by: Bantel, Linus, et al.
Published: (2026)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
by: Cho, Jaehong, et al.
Published: (2024)
by: Cho, Jaehong, et al.
Published: (2024)
Unlocking the Edge deployment and ondevice acceleration of multi-LoRA enabled one-for-all foundational LLM
by: Kodavanti, Sravanth, et al.
Published: (2026)
by: Kodavanti, Sravanth, et al.
Published: (2026)
InstCache: A Predictive Cache for LLM Serving
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023)
by: Chen, Aodong, et al.
Published: (2023)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
by: Li, Allison, et al.
Published: (2025)
by: Li, Allison, et al.
Published: (2025)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
by: Zhao, Zhan, et al.
Published: (2026)
by: Zhao, Zhan, et al.
Published: (2026)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
by: Wei, Anjiang, et al.
Published: (2024)
by: Wei, Anjiang, et al.
Published: (2024)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
Similar Items
-
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
by: Li, Weiqing, et al.
Published: (2025) -
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
by: Zhao, Bingzhe, et al.
Published: (2025) -
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
by: Jiang, Bo, et al.
Published: (2025) -
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
by: Zhang, Huawei, et al.
Published: (2025) -
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
by: Zhu, Yue, et al.
Published: (2025)