MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xuan, Mo, yue, Zhang, Weigang, Wu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918137388072960
author Xuan, Mo
yue, Zhang
Weigang, Wu
author_facet Xuan, Mo
yue, Zhang
Weigang, Wu
contents Model-as-a-Service (MaaS) platforms face diverse Service Level Objective (SLO) requirements stemming from various large language model (LLM) applications, manifested in contextual complexity, first-token latency, and between-token latency. On the other hand, an LLM instance, when configured with different parallelism strategies and inference batch sizes, exhibits distinct performance characteristics and can thus be used to serve different SLO requirements. However, current LLM inference systems typically deploy instances of the same model with identical configurations, lacking mechanisms to leverage such heterogeneity. To fill this research gap, we propose MaaSO, the first MaaS Orchestrator, which comprises three modules: (1) a profiler characterizing instance performance under diverse parallelism strategies and inference batch sizes; (2) a placer optimizing heterogeneous instance configurations; (3) a distributor enabling SLO-aware request distribution and preventing cascaded timeouts in continuous batching. Experiments show that MaaSO improves the SLO satisfaction ratio by 15 to 30% and reduces response latency by 40 to 60% compared to existing approaches, and significantly lowers overall orchestration overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06362
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
Xuan, Mo
yue, Zhang
Weigang, Wu
Distributed, Parallel, and Cluster Computing
Model-as-a-Service (MaaS) platforms face diverse Service Level Objective (SLO) requirements stemming from various large language model (LLM) applications, manifested in contextual complexity, first-token latency, and between-token latency. On the other hand, an LLM instance, when configured with different parallelism strategies and inference batch sizes, exhibits distinct performance characteristics and can thus be used to serve different SLO requirements. However, current LLM inference systems typically deploy instances of the same model with identical configurations, lacking mechanisms to leverage such heterogeneity. To fill this research gap, we propose MaaSO, the first MaaS Orchestrator, which comprises three modules: (1) a profiler characterizing instance performance under diverse parallelism strategies and inference batch sizes; (2) a placer optimizing heterogeneous instance configurations; (3) a distributor enabling SLO-aware request distribution and preventing cascaded timeouts in continuous batching. Experiments show that MaaSO improves the SLO satisfaction ratio by 15 to 30% and reduces response latency by 40 to 60% compared to existing approaches, and significantly lowers overall orchestration overhead.
title MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.06362