HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sun, Ting, Wang, Penghan, Lai, Fan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918178004664320
author Sun, Ting
Wang, Penghan
Lai, Fan
author_facet Sun, Ting
Wang, Penghan
Lai, Fan
contents Large language models (LLMs) have facilitated a wide range of applications with distinct service-level objectives (SLOs), from latency-sensitive online tasks like interactive chatbots to throughput-oriented offline workloads like data synthesis. The existing deployment model, which dedicates machines to each workload, simplifies SLO management but often leads to poor resource utilization. This paper introduces HyGen, an interference-aware LLM serving system that enables efficient co-location of online and offline workloads while preserving SLOs. HyGen incorporates two key innovations: (1) performance control mechanisms, including a latency predictor to estimate batch execution time and an SLO-aware profiler to quantify latency interference, and (2) SLO-aware offline scheduling policies that maximize serving throughput and prevent starvation. Our evaluation on production workloads shows that HyGen achieves up to 3.9-5.8x throughput gains over online and hybrid serving baselines, while ensuring latency SLOs. The code of HyGen is publicly available at https://github.com/UIUC-MLSys/HyGen.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
Sun, Ting
Wang, Penghan
Lai, Fan
Distributed, Parallel, and Cluster Computing
Machine Learning
Large language models (LLMs) have facilitated a wide range of applications with distinct service-level objectives (SLOs), from latency-sensitive online tasks like interactive chatbots to throughput-oriented offline workloads like data synthesis. The existing deployment model, which dedicates machines to each workload, simplifies SLO management but often leads to poor resource utilization. This paper introduces HyGen, an interference-aware LLM serving system that enables efficient co-location of online and offline workloads while preserving SLOs. HyGen incorporates two key innovations: (1) performance control mechanisms, including a latency predictor to estimate batch execution time and an SLO-aware profiler to quantify latency interference, and (2) SLO-aware offline scheduling policies that maximize serving throughput and prevent starvation. Our evaluation on production workloads shows that HyGen achieves up to 3.9-5.8x throughput gains over online and hybrid serving baselines, while ensuring latency SLOs. The code of HyGen is publicly available at https://github.com/UIUC-MLSys/HyGen.
title HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2501.14808