MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Cunchen, Huang, Heyang, Hu, Junhao, Xu, Jiang, Chen, Xusheng, Xie, Tao, Wang, Chenxi, Wang, Sa, Bao, Yungang, Sun, Ninghui, Shan, Yizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910757768134656
author Hu, Cunchen
Huang, Heyang
Hu, Junhao
Xu, Jiang
Chen, Xusheng
Xie, Tao
Wang, Chenxi
Wang, Sa
Bao, Yungang
Sun, Ninghui
Shan, Yizhou
author_facet Hu, Cunchen
Huang, Heyang
Hu, Junhao
Xu, Jiang
Chen, Xusheng
Xie, Tao
Wang, Chenxi
Wang, Sa
Bao, Yungang
Sun, Ninghui
Shan, Yizhou
contents Large language model (LLM) serving has transformed from stateless to stateful systems, utilizing techniques like context caching and disaggregated inference. These optimizations extend the lifespan and domain of the KV cache, necessitating a new architectural approach. We present MemServe, a unified system that integrates both inter-request and intra-request optimizations. MemServe introduces MemPool, an elastic memory pool managing distributed memory and KV caches across serving instances. Using MemPool APIs, MemServe combines context caching with disaggregated inference for the first time, supported by a global scheduler that enhances cache reuse through a global prompt tree-based locality-aware policy. Tests show that MemServe significantly improves job completion time and time-to-first-time.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17565
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Hu, Cunchen
Huang, Heyang
Hu, Junhao
Xu, Jiang
Chen, Xusheng
Xie, Tao
Wang, Chenxi
Wang, Sa
Bao, Yungang
Sun, Ninghui
Shan, Yizhou
Distributed, Parallel, and Cluster Computing
Large language model (LLM) serving has transformed from stateless to stateful systems, utilizing techniques like context caching and disaggregated inference. These optimizations extend the lifespan and domain of the KV cache, necessitating a new architectural approach. We present MemServe, a unified system that integrates both inter-request and intra-request optimizations. MemServe introduces MemPool, an elastic memory pool managing distributed memory and KV caches across serving instances. Using MemPool APIs, MemServe combines context caching with disaggregated inference for the first time, supported by a global scheduler that enhances cache reuse through a global prompt tree-based locality-aware policy. Tests show that MemServe significantly improves job completion time and time-to-first-time.
title MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2406.17565