Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Yue, Yu, Hao, Wang, Chen, Liu, Zhuoran, Lee, Eun Kyung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915308908838912
author Zhu, Yue
Yu, Hao
Wang, Chen
Liu, Zhuoran
Lee, Eun Kyung
author_facet Zhu, Yue
Yu, Hao
Wang, Chen
Liu, Zhuoran
Lee, Eun Kyung
contents The increasing adoption of large language models (LLMs) with extended context windows necessitates efficient Key-Value Cache (KVC) management to optimize inference performance. Inference workloads like Retrieval-Augmented Generation (RAG) and agents exhibit high cache reusability, making efficient caching critical to reducing redundancy and improving speed. We analyze real-world KVC access patterns using publicly available traces and evaluate commercial key-value stores like Redis and state-of-the-art RDMA-based systems (CHIME [1] and Sherman [2]) for KVC metadata management. Our work demonstrates the lack of tailored storage solution for KVC prefilling, underscores the need for an efficient distributed caching system with optimized metadata management for LLM workloads, and provides insights into designing improved KVC management systems for scalable, low-latency inference.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
Zhu, Yue
Yu, Hao
Wang, Chen
Liu, Zhuoran
Lee, Eun Kyung
Emerging Technologies
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
The increasing adoption of large language models (LLMs) with extended context windows necessitates efficient Key-Value Cache (KVC) management to optimize inference performance. Inference workloads like Retrieval-Augmented Generation (RAG) and agents exhibit high cache reusability, making efficient caching critical to reducing redundancy and improving speed. We analyze real-world KVC access patterns using publicly available traces and evaluate commercial key-value stores like Redis and state-of-the-art RDMA-based systems (CHIME [1] and Sherman [2]) for KVC metadata management. Our work demonstrates the lack of tailored storage solution for KVC prefilling, underscores the need for an efficient distributed caching system with optimized metadata management for LLM workloads, and provides insights into designing improved KVC management systems for scalable, low-latency inference.
title Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
topic Emerging Technologies
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2505.21919