BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908396128567296 |
|---|---|
| author | Hu, Xiannan Zeng, Tianyou Yuan, Xiaoming Song, Liwei Zhang, Guangyuan He, Bangzheng |
| author_facet | Hu, Xiannan Zeng, Tianyou Yuan, Xiaoming Song, Liwei Zhang, Guangyuan He, Bangzheng |
| contents | Serving large language models (LLMs) to millions of users requires efficient resource allocation and parallelism strategies. It is a labor intensive trial-and-error process to find such a strategy. We present BestServe, a novel framework for ranking serving strategies by estimating goodput under various operating scenarios. Supporting both collocated and disaggregated architectures, BestServe leverages an inference simulator built on an adapted roofline model and CPU-GPU dispatch dynamics. Our framework determines the optimal strategy in minutes on a single standard CPU, eliminating the need for costly benchmarking, while achieving predictions within a $20\%$ error margin. It appeals to be practical for rapid deployment planning because of its lightweight design and strong extensibility. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_05871 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Hu, Xiannan Zeng, Tianyou Yuan, Xiaoming Song, Liwei Zhang, Guangyuan He, Bangzheng Machine Learning Distributed, Parallel, and Cluster Computing Performance Serving large language models (LLMs) to millions of users requires efficient resource allocation and parallelism strategies. It is a labor intensive trial-and-error process to find such a strategy. We present BestServe, a novel framework for ranking serving strategies by estimating goodput under various operating scenarios. Supporting both collocated and disaggregated architectures, BestServe leverages an inference simulator built on an adapted roofline model and CPU-GPU dispatch dynamics. Our framework determines the optimal strategy in minutes on a single standard CPU, eliminating the need for costly benchmarking, while achieving predictions within a $20\%$ error margin. It appeals to be practical for rapid deployment planning because of its lightweight design and strong extensibility. |
| title | BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing Performance |
| url | https://arxiv.org/abs/2506.05871 |