Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908515840294912 |
|---|---|
| author | Qin, Ruoyu Li, Zheming He, Weiran Zhang, Mingxing Wu, Yongwei Zheng, Weimin Xu, Xinran |
| author_facet | Qin, Ruoyu Li, Zheming He, Weiran Zhang, Mingxing Wu, Yongwei Zheng, Weimin Xu, Xinran |
| contents | Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_00079 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving Qin, Ruoyu Li, Zheming He, Weiran Zhang, Mingxing Wu, Yongwei Zheng, Weimin Xu, Xinran Distributed, Parallel, and Cluster Computing Artificial Intelligence Hardware Architecture Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests. |
| title | Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence Hardware Architecture |
| url | https://arxiv.org/abs/2407.00079 |