OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Siyu, Tang, Zihan, Zeng, Yuting, Chen, Hui, Ding, Guiguang, Liu, Tongxuan, Zhang, Ke, Yang, Hailong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917107851067392
author Wu, Siyu
Tang, Zihan
Zeng, Yuting
Chen, Hui
Ding, Guiguang
Liu, Tongxuan
Zhang, Ke
Yang, Hailong
author_facet Wu, Siyu
Tang, Zihan
Zeng, Yuting
Chen, Hui
Ding, Guiguang
Liu, Tongxuan
Zhang, Ke
Yang, Hailong
contents Large Language Models (LLMs) are increasingly deployed in both latency-sensitive online services and cost-sensitive offline workloads. Co-locating these workloads on shared serving instances can improve resource utilization, but directly applying this approach to Prefill/Decode (P/D) disaggregated systems introduces severe load imbalance, as fluctuating request mixes alter the intrinsic P/D ratio. Existing dynamic adjustment techniques cannot keep up with the bursty traffic patterns of online services. We propose a latency-constraint disaggregated architecture, which separates cluster resources into latency-strict and latency-relaxed pools based on task latency requirements. This design enables flexible placement of offline decode tasks, mitigating P/D imbalance while preserving online performance. To fully exploit this flexibility, we propose (1) a bottleneck-based scheduler guided by a Roofline-based performance model for performance bottleneck based scheduling, and (2) a fast preemption mechanism that strictly enforces Service Level Objectives (SLOs) for online requests. Experiments on real-world traces show that compared to existing offline system approaches, our method improves offline throughput by up to 3x, while maintaining online request SLOs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
Wu, Siyu
Tang, Zihan
Zeng, Yuting
Chen, Hui
Ding, Guiguang
Liu, Tongxuan
Zhang, Ke
Yang, Hailong
Distributed, Parallel, and Cluster Computing
Large Language Models (LLMs) are increasingly deployed in both latency-sensitive online services and cost-sensitive offline workloads. Co-locating these workloads on shared serving instances can improve resource utilization, but directly applying this approach to Prefill/Decode (P/D) disaggregated systems introduces severe load imbalance, as fluctuating request mixes alter the intrinsic P/D ratio. Existing dynamic adjustment techniques cannot keep up with the bursty traffic patterns of online services. We propose a latency-constraint disaggregated architecture, which separates cluster resources into latency-strict and latency-relaxed pools based on task latency requirements. This design enables flexible placement of offline decode tasks, mitigating P/D imbalance while preserving online performance. To fully exploit this flexibility, we propose (1) a bottleneck-based scheduler guided by a Roofline-based performance model for performance bottleneck based scheduling, and (2) a fast preemption mechanism that strictly enforces Service Level Objectives (SLOs) for online requests. Experiments on real-world traces show that compared to existing offline system approaches, our method improves offline throughput by up to 3x, while maintaining online request SLOs.
title OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2511.21862