DeepServe: Serverless Large Language Model Serving at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Junhao, Xu, Jiang, Liu, Zhixia, He, Yulong, Chen, Yuetao, Xu, Hao, Liu, Jiang, Meng, Jie, Zhang, Baoquan, Wan, Shining, Dan, Gengyuan, Dong, Zhiyu, Ren, Zhihao, Liu, Changhong, Xie, Tao, Lin, Dayun, Zhang, Qin, Yu, Yue, Feng, Hao, Chen, Xusheng, Shan, Yizhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912418572009472
author Hu, Junhao
Xu, Jiang
Liu, Zhixia
He, Yulong
Chen, Yuetao
Xu, Hao
Liu, Jiang
Meng, Jie
Zhang, Baoquan
Wan, Shining
Dan, Gengyuan
Dong, Zhiyu
Ren, Zhihao
Liu, Changhong
Xie, Tao
Lin, Dayun
Zhang, Qin
Yu, Yue
Feng, Hao
Chen, Xusheng
Shan, Yizhou
author_facet Hu, Junhao
Xu, Jiang
Liu, Zhixia
He, Yulong
Chen, Yuetao
Xu, Hao
Liu, Jiang
Meng, Jie
Zhang, Baoquan
Wan, Shining
Dan, Gengyuan
Dong, Zhiyu
Ren, Zhihao
Liu, Changhong
Xie, Tao
Lin, Dayun
Zhang, Qin
Yu, Yue
Feng, Hao
Chen, Xusheng
Shan, Yizhou
contents In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across posttraining and model-serving tasks. Second, DEEPSERVE integrates an in-house serving engine named FLOWSERVE using a microkernel-inspired design, NPU-centric execution, and SPMD-based parallelism to optimize LLM serving. Third, DEEPSERVE includes novel scheduling policies tailored for a configuration with both PD-disaggregated and PD-colocated instances. Fourth, DEEPSERVE includes optimizations such as pre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in seconds. DEEPSERVE has been in production for over a year, operating on a large Ascend NPU cluster and providing industrystandard APIs for fine-tuning, agent serving, and model serving to our customers.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepServe: Serverless Large Language Model Serving at Scale
Hu, Junhao
Xu, Jiang
Liu, Zhixia
He, Yulong
Chen, Yuetao
Xu, Hao
Liu, Jiang
Meng, Jie
Zhang, Baoquan
Wan, Shining
Dan, Gengyuan
Dong, Zhiyu
Ren, Zhihao
Liu, Changhong
Xie, Tao
Lin, Dayun
Zhang, Qin
Yu, Yue
Feng, Hao
Chen, Xusheng
Shan, Yizhou
Distributed, Parallel, and Cluster Computing
In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across posttraining and model-serving tasks. Second, DEEPSERVE integrates an in-house serving engine named FLOWSERVE using a microkernel-inspired design, NPU-centric execution, and SPMD-based parallelism to optimize LLM serving. Third, DEEPSERVE includes novel scheduling policies tailored for a configuration with both PD-disaggregated and PD-colocated instances. Fourth, DEEPSERVE includes optimizations such as pre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in seconds. DEEPSERVE has been in production for over a year, operating on a large Ascend NPU cluster and providing industrystandard APIs for fine-tuning, agent serving, and model serving to our customers.
title DeepServe: Serverless Large Language Model Serving at Scale
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2501.14417