ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Kaiwen, Tan, Xin, Yu, Minchen, Li, Jingzong, Xu, Hong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909040734371840
author Chen, Kaiwen
Tan, Xin
Yu, Minchen
Li, Jingzong
Xu, Hong
author_facet Chen, Kaiwen
Tan, Xin
Yu, Minchen
Li, Jingzong
Xu, Hong
contents Large Reasoning Models (LRMs) are becoming integral to many AI inference systems, enhancing their capabilities with advanced reasoning. However, deploying these models in production environments presents a significant QoS challenge: the substantial memory overhead from their long, auto-regressive inference processes severely limits throughput and increases latency, thereby affecting the quality of service for concurrent users. We observe that LRMs frequently generate highly similar intermediate reasoning steps, which, in turn, correspond to highly similar KV cache states across layers. Building on this insight, we propose ReasonCache, a novel KV cache management approach designed to improve the QoS of AI inference systems. ReasonCache utilizes a Collaborative Filtering Algorithm to efficiently identify reusable KV cache blocks and enables zero-copy cache reuse. Experimental evaluation demonstrates that ReasonCache achieves a peak throughput improvement of 89.2% and an average gain of 40-60%, leading to more responsive and cost-effective AI inference services. Notably, this performance is achieved while maintaining higher accuracy compared to existing KV cache management techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21433
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
Chen, Kaiwen
Tan, Xin
Yu, Minchen
Li, Jingzong
Xu, Hong
Machine Learning
Artificial Intelligence
Large Reasoning Models (LRMs) are becoming integral to many AI inference systems, enhancing their capabilities with advanced reasoning. However, deploying these models in production environments presents a significant QoS challenge: the substantial memory overhead from their long, auto-regressive inference processes severely limits throughput and increases latency, thereby affecting the quality of service for concurrent users. We observe that LRMs frequently generate highly similar intermediate reasoning steps, which, in turn, correspond to highly similar KV cache states across layers. Building on this insight, we propose ReasonCache, a novel KV cache management approach designed to improve the QoS of AI inference systems. ReasonCache utilizes a Collaborative Filtering Algorithm to efficiently identify reusable KV cache blocks and enables zero-copy cache reuse. Experimental evaluation demonstrates that ReasonCache achieves a peak throughput improvement of 89.2% and an average gain of 40-60%, leading to more responsive and cost-effective AI inference services. Notably, this performance is achieved while maintaining higher accuracy compared to existing KV cache management techniques.
title ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.21433