Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Rongzhi, Du, Ruogu, Chu, Zefang, Zhao, Sida, Han, Chunlei, Shi, Zuocheng, Shao, Yiwen, Han, Huanle, Huang, Long, Liu, Zherui, Liu, Shufan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912556173492224
author Li, Rongzhi
Du, Ruogu
Chu, Zefang
Zhao, Sida
Han, Chunlei
Shi, Zuocheng
Shao, Yiwen
Han, Huanle
Huang, Long
Liu, Zherui
Liu, Shufan
author_facet Li, Rongzhi
Du, Ruogu
Chu, Zefang
Zhao, Sida
Han, Chunlei
Shi, Zuocheng
Shao, Yiwen
Han, Huanle
Huang, Long
Liu, Zherui
Liu, Shufan
contents Serving Large Language Models (LLMs) is a GPU-intensive task where traditional autoscalers fall short, particularly for modern Prefill-Decode (P/D) disaggregated architectures. This architectural shift, while powerful, introduces significant operational challenges, including inefficient use of heterogeneous hardware, network bottlenecks, and critical imbalances between prefill and decode stages. We introduce HeteroScale, a coordinated autoscaling framework that addresses the core challenges of P/D disaggregated serving. HeteroScale combines a topology-aware scheduler that adapts to heterogeneous hardware and network constraints with a novel metric-driven policy derived from the first large-scale empirical study of autoscaling signals in production. By leveraging a single, robust metric to jointly scale prefill and decode pools, HeteroScale maintains architectural balance while ensuring efficient, adaptive resource management. Deployed in a massive production environment on tens of thousands of GPUs, HeteroScale has proven its effectiveness, increasing average GPU utilization by a significant 26.6 percentage points and saving hundreds of thousands of GPU-hours daily, all while upholding stringent service level objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
Li, Rongzhi
Du, Ruogu
Chu, Zefang
Zhao, Sida
Han, Chunlei
Shi, Zuocheng
Shao, Yiwen
Han, Huanle
Huang, Long
Liu, Zherui
Liu, Shufan
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Serving Large Language Models (LLMs) is a GPU-intensive task where traditional autoscalers fall short, particularly for modern Prefill-Decode (P/D) disaggregated architectures. This architectural shift, while powerful, introduces significant operational challenges, including inefficient use of heterogeneous hardware, network bottlenecks, and critical imbalances between prefill and decode stages. We introduce HeteroScale, a coordinated autoscaling framework that addresses the core challenges of P/D disaggregated serving. HeteroScale combines a topology-aware scheduler that adapts to heterogeneous hardware and network constraints with a novel metric-driven policy derived from the first large-scale empirical study of autoscaling signals in production. By leveraging a single, robust metric to jointly scale prefill and decode pools, HeteroScale maintains architectural balance while ensuring efficient, adaptive resource management. Deployed in a massive production environment on tens of thousands of GPUs, HeteroScale has proven its effectiveness, increasing average GPU utilization by a significant 26.6 percentage points and saving hundreds of thousands of GPU-hours daily, all while upholding stringent service level objectives.
title Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2508.19559