Comparative Characterization of KV Cache Management Strategies for LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mamo, Oteo, Kogiou, Olga, Yi, Hyunjin, Yu, Weikuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
von: Kim, Dowon, et al.
Veröffentlicht: (2025)
von: Kim, Dowon, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
von: Xia, Tianhua, et al.
Veröffentlicht: (2025)
von: Xia, Tianhua, et al.
Veröffentlicht: (2025)
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS
von: Rajesh, Varun, et al.
Veröffentlicht: (2025)
von: Rajesh, Varun, et al.
Veröffentlicht: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
von: Yao, Jiayi, et al.
Veröffentlicht: (2026)
von: Yao, Jiayi, et al.
Veröffentlicht: (2026)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
von: Yun, Sungmin, et al.
Veröffentlicht: (2025)
von: Yun, Sungmin, et al.
Veröffentlicht: (2025)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025)
von: Du, Dayou, et al.
Veröffentlicht: (2025)
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
von: Li, Pu, et al.
Veröffentlicht: (2026)
von: Li, Pu, et al.
Veröffentlicht: (2026)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
von: He, Zifan, et al.
Veröffentlicht: (2025)
von: He, Zifan, et al.
Veröffentlicht: (2025)
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
von: Zou, Guoqiang, et al.
Veröffentlicht: (2025)
von: Zou, Guoqiang, et al.
Veröffentlicht: (2025)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024)
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
von: You, Kang, et al.
Veröffentlicht: (2026)
von: You, Kang, et al.
Veröffentlicht: (2026)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
von: Javat, Abdurrahman, et al.
Veröffentlicht: (2026)
von: Javat, Abdurrahman, et al.
Veröffentlicht: (2026)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
von: Xu, Weikai, et al.
Veröffentlicht: (2025)
von: Xu, Weikai, et al.
Veröffentlicht: (2025)
Efficient LLM inference solution on Intel GPU
von: Wu, Hui, et al.
Veröffentlicht: (2023)
von: Wu, Hui, et al.
Veröffentlicht: (2023)
PrefixLLM: LLM-aided Prefix Circuit Design
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
LLM-DSE: Searching Accelerator Parameters with LLM Agents
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
von: Zhu, Zhantong, et al.
Veröffentlicht: (2025)
von: Zhu, Zhantong, et al.
Veröffentlicht: (2025)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
LLM Inference Acceleration via Efficient Operation Fusion
von: Salmani, Mahsa, et al.
Veröffentlicht: (2025)
von: Salmani, Mahsa, et al.
Veröffentlicht: (2025)
CacheMind: From Miss Rates to Why -- Natural-Language, Trace-Grounded Reasoning for Cache Replacement
von: Mhapsekar, Kaushal, et al.
Veröffentlicht: (2026)
von: Mhapsekar, Kaushal, et al.
Veröffentlicht: (2026)
AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications
von: Wu, Yuchao, et al.
Veröffentlicht: (2026)
von: Wu, Yuchao, et al.
Veröffentlicht: (2026)
Surrogates, Spikes, and Sparsity: Performance Analysis and Characterization of SNN Hyperparameters on Hardware
von: Aliyev, Ilkin, et al.
Veröffentlicht: (2026)
von: Aliyev, Ilkin, et al.
Veröffentlicht: (2026)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
von: Errabii, Sohaib, et al.
Veröffentlicht: (2026)
von: Errabii, Sohaib, et al.
Veröffentlicht: (2026)
ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
von: Lou, Binglei, et al.
Veröffentlicht: (2026)
von: Lou, Binglei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
von: Kim, Dowon, et al.
Veröffentlicht: (2025) -
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
von: Fang, Yunhua, et al.
Veröffentlicht: (2025) -
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
von: Xia, Tianhua, et al.
Veröffentlicht: (2025) -
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025) -
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)