Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yunzhao, Xu, Qiang, Hu, Y. Charlie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914050463498240
author Liu, Yunzhao
Xu, Qiang
Hu, Y. Charlie
author_facet Liu, Yunzhao
Xu, Qiang
Hu, Y. Charlie
contents Efficient LLM inference is critical for real-world applications, especially within heterogeneous GPU clusters commonly found in organizations and on-premise datacenters as GPU architecture rapidly evolves. Current disaggregated prefill strategies, which separate the prefill and decode stages of LLM inference across different GPUs, often suffer from suboptimal performance due to imbalances between GPU capabilities and workload demands. On the other hand, extending conventional data parallelism and pipeline parallelism to heterogeneous setups incurs high inference latencies. To address these challenges, we introduce Cronus, a novel LLM inference system designed to dynamically balance workloads across heterogeneous GPUs using partially disaggregated prefill. Cronus partitions each prefill stage and executes its initial portion on the low-end GPU, while overlapping the remaining prefill and decode stages of earlier requests on the high-end GPU. Extensive evaluations across various high-end and low-end GPU combinations demonstrate that Cronus significantly improves the throughput over disaggregated prefill. It also reduces TTFT P99 and TBT P99 significantly over DP and PP while maintaining similar or better throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
Liu, Yunzhao
Xu, Qiang
Hu, Y. Charlie
Distributed, Parallel, and Cluster Computing
Efficient LLM inference is critical for real-world applications, especially within heterogeneous GPU clusters commonly found in organizations and on-premise datacenters as GPU architecture rapidly evolves. Current disaggregated prefill strategies, which separate the prefill and decode stages of LLM inference across different GPUs, often suffer from suboptimal performance due to imbalances between GPU capabilities and workload demands. On the other hand, extending conventional data parallelism and pipeline parallelism to heterogeneous setups incurs high inference latencies. To address these challenges, we introduce Cronus, a novel LLM inference system designed to dynamically balance workloads across heterogeneous GPUs using partially disaggregated prefill. Cronus partitions each prefill stage and executes its initial portion on the low-end GPU, while overlapping the remaining prefill and decode stages of earlier requests on the high-end GPU. Extensive evaluations across various high-end and low-end GPU combinations demonstrate that Cronus significantly improves the throughput over disaggregated prefill. It also reduces TTFT P99 and TBT P99 significantly over DP and PP while maintaining similar or better throughput.
title Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.17357