MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910757768134656 |
|---|---|
| author | Hu, Cunchen Huang, Heyang Hu, Junhao Xu, Jiang Chen, Xusheng Xie, Tao Wang, Chenxi Wang, Sa Bao, Yungang Sun, Ninghui Shan, Yizhou |
| author_facet | Hu, Cunchen Huang, Heyang Hu, Junhao Xu, Jiang Chen, Xusheng Xie, Tao Wang, Chenxi Wang, Sa Bao, Yungang Sun, Ninghui Shan, Yizhou |
| contents | Large language model (LLM) serving has transformed from stateless to stateful systems, utilizing techniques like context caching and disaggregated inference. These optimizations extend the lifespan and domain of the KV cache, necessitating a new architectural approach. We present MemServe, a unified system that integrates both inter-request and intra-request optimizations. MemServe introduces MemPool, an elastic memory pool managing distributed memory and KV caches across serving instances. Using MemPool APIs, MemServe combines context caching with disaggregated inference for the first time, supported by a global scheduler that enhances cache reuse through a global prompt tree-based locality-aware policy. Tests show that MemServe significantly improves job completion time and time-to-first-time. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_17565 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool Hu, Cunchen Huang, Heyang Hu, Junhao Xu, Jiang Chen, Xusheng Xie, Tao Wang, Chenxi Wang, Sa Bao, Yungang Sun, Ninghui Shan, Yizhou Distributed, Parallel, and Cluster Computing Large language model (LLM) serving has transformed from stateless to stateful systems, utilizing techniques like context caching and disaggregated inference. These optimizations extend the lifespan and domain of the KV cache, necessitating a new architectural approach. We present MemServe, a unified system that integrates both inter-request and intra-request optimizations. MemServe introduces MemPool, an elastic memory pool managing distributed memory and KV caches across serving instances. Using MemPool APIs, MemServe combines context caching with disaggregated inference for the first time, supported by a global scheduler that enhances cache reuse through a global prompt tree-based locality-aware policy. Tests show that MemServe significantly improves job completion time and time-to-first-time. |
| title | MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2406.17565 |