Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Ruoyu, Li, Zheming, He, Weiran, Zhang, Mingxing, Wu, Yongwei, Zheng, Weimin, Xu, Xinran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908515840294912
author Qin, Ruoyu
Li, Zheming
He, Weiran
Zhang, Mingxing
Wu, Yongwei
Zheng, Weimin
Xu, Xinran
author_facet Qin, Ruoyu
Li, Zheming
He, Weiran
Zhang, Mingxing
Wu, Yongwei
Zheng, Weimin
Xu, Xinran
contents Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00079
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Qin, Ruoyu
Li, Zheming
He, Weiran
Zhang, Mingxing
Wu, Yongwei
Zheng, Weimin
Xu, Xinran
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hardware Architecture
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.
title Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2407.00079