Compute Or Load KV Cache? Why Not Both?
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Shuowei, Liu, Xueshen, Zhang, Qingzhao, Mao, Z. Morley |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eagle: Efficient Training-Free Router for Multi-LLM Inference
by: Zhao, Zesen, et al.
Published: (2024)
by: Zhao, Zesen, et al.
Published: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
by: Hu, Pingbang, et al.
Published: (2026)
by: Hu, Pingbang, et al.
Published: (2026)
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
by: Zheng, Haizhong, et al.
Published: (2026)
by: Zheng, Haizhong, et al.
Published: (2026)
AutoSpec: Automated Generation of Neural Network Specifications
by: Jin, Shuowei, et al.
Published: (2024)
by: Jin, Shuowei, et al.
Published: (2024)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
by: Liu, Xueshen, et al.
Published: (2026)
by: Liu, Xueshen, et al.
Published: (2026)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
by: Zheng, Haizhong, et al.
Published: (2024)
by: Zheng, Haizhong, et al.
Published: (2024)
MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search
by: Cho, Minkyoung, et al.
Published: (2026)
by: Cho, Minkyoung, et al.
Published: (2026)
Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
Plato: Plan to Efficiently Decode for Large Language Model Inference
by: Jin, Shuowei, et al.
Published: (2024)
by: Jin, Shuowei, et al.
Published: (2024)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
by: Sun, Qiheng, et al.
Published: (2025)
by: Sun, Qiheng, et al.
Published: (2025)
KVSculpt: KV Cache Compression as Distillation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
HSTFL: A Heterogeneous Federated Learning Framework for Misaligned Spatiotemporal Forecasting
by: Cai, Shuowei, et al.
Published: (2024)
by: Cai, Shuowei, et al.
Published: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving
by: Zhu, Ruiyang, et al.
Published: (2026)
by: Zhu, Ruiyang, et al.
Published: (2026)
Sparse Attention across Multiple-context KV Cache
by: Cao, Ziyi, et al.
Published: (2025)
by: Cao, Ziyi, et al.
Published: (2025)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
by: Roy, Sourjya, et al.
Published: (2025)
by: Roy, Sourjya, et al.
Published: (2025)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
by: Chen, Chuangtao, et al.
Published: (2026)
by: Chen, Chuangtao, et al.
Published: (2026)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
by: Geng, Yingsheng, et al.
Published: (2026)
by: Geng, Yingsheng, et al.
Published: (2026)
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
by: Tan, Yifan, et al.
Published: (2024)
by: Tan, Yifan, et al.
Published: (2024)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
by: Wu, Wenbo, et al.
Published: (2025)
by: Wu, Wenbo, et al.
Published: (2025)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
by: Mao, Yuzhen, et al.
Published: (2026)
by: Mao, Yuzhen, et al.
Published: (2026)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
by: Yang, Bin, et al.
Published: (2025)
by: Yang, Bin, et al.
Published: (2025)
Cocoon: Robust Multi-Modal Perception with Uncertainty-Aware Sensor Fusion
by: Cho, Minkyoung, et al.
Published: (2024)
by: Cho, Minkyoung, et al.
Published: (2024)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
by: Liu, Sihao, et al.
Published: (2026)
by: Liu, Sihao, et al.
Published: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
by: Liu, Akide, et al.
Published: (2024)
by: Liu, Akide, et al.
Published: (2024)
FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management
by: Takbir, Nazmul, et al.
Published: (2025)
by: Takbir, Nazmul, et al.
Published: (2025)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
by: Gokhale, Sai, et al.
Published: (2025)
by: Gokhale, Sai, et al.
Published: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
Similar Items
-
Eagle: Efficient Training-Free Router for Multi-LLM Inference
by: Zhao, Zesen, et al.
Published: (2024) -
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
by: Wu, Yongji, et al.
Published: (2025) -
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
by: Hu, Pingbang, et al.
Published: (2026) -
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
by: Zheng, Haizhong, et al.
Published: (2026) -
AutoSpec: Automated Generation of Neural Network Specifications
by: Jin, Shuowei, et al.
Published: (2024)