Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Da, Wei, Kalyvianaki, Evangelia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909735963328512
author Da, Wei
Kalyvianaki, Evangelia
author_facet Da, Wei
Kalyvianaki, Evangelia
contents This paper presents Block, a distributed scheduling framework designed to optimize load balancing and auto-provisioning across instances in large language model serving frameworks by leveraging contextual information from incoming requests. Unlike popular model serving systems that rely on monolithic and heuristic task schedulers, Block operates as a fully distributed, stateless, and predictive scheduling system to achieve low overhead, reliability, and scalability. It leverages the deterministic and predictable characteristics of LLM inferences, such as host configurations, response lengths, and hardware performance, to make scheduling decisions based on accurately predicted metrics. Evaluation on a 12 GPUs cluster shows that Block significantly outperforms heuristic schedulers, boosting serving capacity by up to 16.7\% and reducing P99 tail latency by up to 49.5\%. These performance gains remain consistent across diverse models, workloads and configurations. Code and data are open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
Da, Wei
Kalyvianaki, Evangelia
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
This paper presents Block, a distributed scheduling framework designed to optimize load balancing and auto-provisioning across instances in large language model serving frameworks by leveraging contextual information from incoming requests. Unlike popular model serving systems that rely on monolithic and heuristic task schedulers, Block operates as a fully distributed, stateless, and predictive scheduling system to achieve low overhead, reliability, and scalability. It leverages the deterministic and predictable characteristics of LLM inferences, such as host configurations, response lengths, and hardware performance, to make scheduling decisions based on accurately predicted metrics. Evaluation on a 12 GPUs cluster shows that Block significantly outperforms heuristic schedulers, boosting serving capacity by up to 16.7\% and reducing P99 tail latency by up to 49.5\%. These performance gains remain consistent across diverse models, workloads and configurations. Code and data are open-sourced.
title Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2508.03611