OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Youhe, Fu, Fangcheng, Wang, Taiyi, He, Guoliang, Yoneki, Eiko
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914526870372352
author Jiang, Youhe
Fu, Fangcheng
Wang, Taiyi
He, Guoliang
Yoneki, Eiko
author_facet Jiang, Youhe
Fu, Fangcheng
Wang, Taiyi
He, Guoliang
Yoneki, Eiko
contents Serving Large Language Models (LLMs) can benefit immensely from parallelizing both the model and input requests across multiple devices, but incoming workloads exhibit substantial spatial and temporal heterogeneity. Spatially, workloads comprise heterogeneous requests with varying compute and memory demands. Temporally, workload composition varies over time. Nevertheless, existing systems typically assume spatially uniform and temporally stable workloads, employing a homogeneous, static model deployment. This mismatch between the assumption and real-world spatial-temporal heterogeneity results in suboptimal performance. We present OServe, an LLM serving system with heterogeneous and flexible model deployment that addresses both spatial and temporal heterogeneity. First, OServe introduces a novel workload-aware scheduling algorithm that optimizes heterogeneous model deployments according to real-time workload characteristics. Second, OServe proposes an efficient workload-adaptive switching method that migrates model deployments in response to predicted workload changes. Experiments on real-world traces show that OServe improves performance by up to 2$\times$ (average: 1.5$\times$) compared to state-of-the-art serving systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12151
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
Jiang, Youhe
Fu, Fangcheng
Wang, Taiyi
He, Guoliang
Yoneki, Eiko
Distributed, Parallel, and Cluster Computing
Serving Large Language Models (LLMs) can benefit immensely from parallelizing both the model and input requests across multiple devices, but incoming workloads exhibit substantial spatial and temporal heterogeneity. Spatially, workloads comprise heterogeneous requests with varying compute and memory demands. Temporally, workload composition varies over time. Nevertheless, existing systems typically assume spatially uniform and temporally stable workloads, employing a homogeneous, static model deployment. This mismatch between the assumption and real-world spatial-temporal heterogeneity results in suboptimal performance. We present OServe, an LLM serving system with heterogeneous and flexible model deployment that addresses both spatial and temporal heterogeneity. First, OServe introduces a novel workload-aware scheduling algorithm that optimizes heterogeneous model deployments according to real-time workload characteristics. Second, OServe proposes an efficient workload-adaptive switching method that migrates model deployments in response to predicted workload changes. Experiments on real-world traces show that OServe improves performance by up to 2$\times$ (average: 1.5$\times$) compared to state-of-the-art serving systems.
title OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.12151