PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Chongpeng, Liao, Xiaojian, Liu, Hancheng, Xiao, Limin, Li, Jianxin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912287190679552
author Liu, Chongpeng
Liao, Xiaojian
Liu, Hancheng
Xiao, Limin
Li, Jianxin
author_facet Liu, Chongpeng
Liao, Xiaojian
Liu, Hancheng
Xiao, Limin
Li, Jianxin
contents This paper presents PipeBoost, a low-latency LLM serving system for multi-GPU (serverless) clusters, which can rapidly launch inference services in response to bursty requests without preemptively over-provisioning GPUs. Many LLM inference tasks rely on the same base model (e.g., LoRA). To leverage this, PipeBoost introduces fault-tolerant pipeline parallelism across both model loading and inference stages. This approach maximizes aggregate PCIe bandwidth and parallel computation across GPUs, enabling faster generation of the first token. PipeBoost also introduces recovery techniques that enable uninterrupted inference services by utilizing the shared advantages of multiple GPUs. Experimental results show that, compared to state-of-the-art low-latency LLM serving systems, PipeBoost reduces inference latency by 31% to 49.8%. For certain models (e.g., OPT-1.3B), PipeBoost achieves cold-start latencies in the range of a few hundred microseconds.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17707
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
Liu, Chongpeng
Liao, Xiaojian
Liu, Hancheng
Xiao, Limin
Li, Jianxin
Distributed, Parallel, and Cluster Computing
This paper presents PipeBoost, a low-latency LLM serving system for multi-GPU (serverless) clusters, which can rapidly launch inference services in response to bursty requests without preemptively over-provisioning GPUs. Many LLM inference tasks rely on the same base model (e.g., LoRA). To leverage this, PipeBoost introduces fault-tolerant pipeline parallelism across both model loading and inference stages. This approach maximizes aggregate PCIe bandwidth and parallel computation across GPUs, enabling faster generation of the first token. PipeBoost also introduces recovery techniques that enable uninterrupted inference services by utilizing the shared advantages of multiple GPUs. Experimental results show that, compared to state-of-the-art low-latency LLM serving systems, PipeBoost reduces inference latency by 31% to 49.8%. For certain models (e.g., OPT-1.3B), PipeBoost achieves cold-start latencies in the range of a few hundred microseconds.
title PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.17707