DeepServe: Serverless Large Language Model Serving at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912418572009472 |
|---|---|
| author | Hu, Junhao Xu, Jiang Liu, Zhixia He, Yulong Chen, Yuetao Xu, Hao Liu, Jiang Meng, Jie Zhang, Baoquan Wan, Shining Dan, Gengyuan Dong, Zhiyu Ren, Zhihao Liu, Changhong Xie, Tao Lin, Dayun Zhang, Qin Yu, Yue Feng, Hao Chen, Xusheng Shan, Yizhou |
| author_facet | Hu, Junhao Xu, Jiang Liu, Zhixia He, Yulong Chen, Yuetao Xu, Hao Liu, Jiang Meng, Jie Zhang, Baoquan Wan, Shining Dan, Gengyuan Dong, Zhiyu Ren, Zhihao Liu, Changhong Xie, Tao Lin, Dayun Zhang, Qin Yu, Yue Feng, Hao Chen, Xusheng Shan, Yizhou |
| contents | In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across posttraining and model-serving tasks. Second, DEEPSERVE integrates an in-house serving engine named FLOWSERVE using a microkernel-inspired design, NPU-centric execution, and SPMD-based parallelism to optimize LLM serving. Third, DEEPSERVE includes novel scheduling policies tailored for a configuration with both PD-disaggregated and PD-colocated instances. Fourth, DEEPSERVE includes optimizations such as pre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in seconds. DEEPSERVE has been in production for over a year, operating on a large Ascend NPU cluster and providing industrystandard APIs for fine-tuning, agent serving, and model serving to our customers. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_14417 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DeepServe: Serverless Large Language Model Serving at Scale Hu, Junhao Xu, Jiang Liu, Zhixia He, Yulong Chen, Yuetao Xu, Hao Liu, Jiang Meng, Jie Zhang, Baoquan Wan, Shining Dan, Gengyuan Dong, Zhiyu Ren, Zhihao Liu, Changhong Xie, Tao Lin, Dayun Zhang, Qin Yu, Yue Feng, Hao Chen, Xusheng Shan, Yizhou Distributed, Parallel, and Cluster Computing In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across posttraining and model-serving tasks. Second, DEEPSERVE integrates an in-house serving engine named FLOWSERVE using a microkernel-inspired design, NPU-centric execution, and SPMD-based parallelism to optimize LLM serving. Third, DEEPSERVE includes novel scheduling policies tailored for a configuration with both PD-disaggregated and PD-colocated instances. Fourth, DEEPSERVE includes optimizations such as pre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in seconds. DEEPSERVE has been in production for over a year, operating on a large Ascend NPU cluster and providing industrystandard APIs for fine-tuning, agent serving, and model serving to our customers. |
| title | DeepServe: Serverless Large Language Model Serving at Scale |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2501.14417 |