GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Du, Boxiao, Huangfu, Boning, Luo, Yizhou, Chen, Chen, Li, Zijun, Yu, Minchen, Fan, Xiaoyi, Guo, Minyi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914572079726592
author Du, Boxiao
Huangfu, Boning
Luo, Yizhou
Chen, Chen
Li, Zijun
Yu, Minchen
Fan, Xiaoyi
Guo, Minyi
author_facet Du, Boxiao
Huangfu, Boning
Luo, Yizhou
Chen, Chen
Li, Zijun
Yu, Minchen
Fan, Xiaoyi
Guo, Minyi
contents Large Language Models (LLMs) play a critical role in emerging agentic applications, where the timely completion of each entire inference is critical. Meanwhile, agentic LLM inferences are increasingly served on heterogeneous GPUs in operator's resource pools. Therefore, it is crucial to route incoming inference requests to appropriate GPUs so that their end-to-end latency requirements are satisfied whenever possible, thereby achieving high goodput. In this paper, we propose GoodServe, a goodput-optimized serving system for agentic inferences over heterogeneous resources. GoodServe performs inference routing in a predict-and-rectify manner. It estimates the request output lengths as well as the GPU serving status in an accurate and also practical manner. Based on information from both the demand and resource sides, it then makes high-quality routing decisions using a just-enough instance selection heuristic. It also periodically monitors SLO-violation risks of active requests and triggers runtime request migrations to address unexpected dynamics. Our evaluations show that GoodServe improves goodput by up to 27.4% over existing routing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16867
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
Du, Boxiao
Huangfu, Boning
Luo, Yizhou
Chen, Chen
Li, Zijun
Yu, Minchen
Fan, Xiaoyi
Guo, Minyi
Distributed, Parallel, and Cluster Computing
Large Language Models (LLMs) play a critical role in emerging agentic applications, where the timely completion of each entire inference is critical. Meanwhile, agentic LLM inferences are increasingly served on heterogeneous GPUs in operator's resource pools. Therefore, it is crucial to route incoming inference requests to appropriate GPUs so that their end-to-end latency requirements are satisfied whenever possible, thereby achieving high goodput. In this paper, we propose GoodServe, a goodput-optimized serving system for agentic inferences over heterogeneous resources. GoodServe performs inference routing in a predict-and-rectify manner. It estimates the request output lengths as well as the GPU serving status in an accurate and also practical manner. Based on information from both the demand and resource sides, it then makes high-quality routing decisions using a just-enough instance selection heuristic. It also periodically monitors SLO-violation risks of active requests and triggers runtime request migrations to address unexpected dynamics. Our evaluations show that GoodServe improves goodput by up to 27.4% over existing routing methods.
title GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.16867