BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xiannan, Zeng, Tianyou, Yuan, Xiaoming, Song, Liwei, Zhang, Guangyuan, He, Bangzheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908396128567296
author Hu, Xiannan
Zeng, Tianyou
Yuan, Xiaoming
Song, Liwei
Zhang, Guangyuan
He, Bangzheng
author_facet Hu, Xiannan
Zeng, Tianyou
Yuan, Xiaoming
Song, Liwei
Zhang, Guangyuan
He, Bangzheng
contents Serving large language models (LLMs) to millions of users requires efficient resource allocation and parallelism strategies. It is a labor intensive trial-and-error process to find such a strategy. We present BestServe, a novel framework for ranking serving strategies by estimating goodput under various operating scenarios. Supporting both collocated and disaggregated architectures, BestServe leverages an inference simulator built on an adapted roofline model and CPU-GPU dispatch dynamics. Our framework determines the optimal strategy in minutes on a single standard CPU, eliminating the need for costly benchmarking, while achieving predictions within a $20\%$ error margin. It appeals to be practical for rapid deployment planning because of its lightweight design and strong extensibility.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05871
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
Hu, Xiannan
Zeng, Tianyou
Yuan, Xiaoming
Song, Liwei
Zhang, Guangyuan
He, Bangzheng
Machine Learning
Distributed, Parallel, and Cluster Computing
Performance
Serving large language models (LLMs) to millions of users requires efficient resource allocation and parallelism strategies. It is a labor intensive trial-and-error process to find such a strategy. We present BestServe, a novel framework for ranking serving strategies by estimating goodput under various operating scenarios. Supporting both collocated and disaggregated architectures, BestServe leverages an inference simulator built on an adapted roofline model and CPU-GPU dispatch dynamics. Our framework determines the optimal strategy in minutes on a single standard CPU, eliminating the need for costly benchmarking, while achieving predictions within a $20\%$ error margin. It appeals to be practical for rapid deployment planning because of its lightweight design and strong extensibility.
title BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2506.05871