Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Yiwei, Chowdhary, Sangeeta, Morris, Nathaniel, Jain, Rutwik, Manne, Srilatha, Bayliss, Sam
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2601.12241
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918294690201600
author Jiang, Yiwei
Chowdhary, Sangeeta
Morris, Nathaniel
Jain, Rutwik
Manne, Srilatha
Bayliss, Sam
author_facet Jiang, Yiwei
Chowdhary, Sangeeta
Morris, Nathaniel
Jain, Rutwik
Manne, Srilatha
Bayliss, Sam
contents Disaggregation has emerged as a powerful strategy for optimizing large language model (LLM) inference by separating compute-intensive prefill and memory-bound decode phases across specialized GPUs. This separation improves utilization and throughput under fixed hardware capacity. However, as model and cluster scales grow, power, rather than compute, has become the dominant limiter of overall performance and cost efficiency. In this paper, we propose RAPID, a power-aware disaggregated inference framework that jointly manages GPU roles and power budgets to sustain goodput within strict power caps. RAPID utilizes static and dynamic power reallocation in addition to GPU reallocation to improve performance under fixed power bounds. RAPID improves overall performance and application consistency beyond what is achievable in current disaggregation solutions, resulting in up to a 2x improvement in SLO attainment at peak load when compared to a static assignment without an increase in complexity or cost.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12241
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Power Aware Dynamic Reallocation For Inference
Jiang, Yiwei
Chowdhary, Sangeeta
Morris, Nathaniel
Jain, Rutwik
Manne, Srilatha
Bayliss, Sam
Distributed, Parallel, and Cluster Computing
Disaggregation has emerged as a powerful strategy for optimizing large language model (LLM) inference by separating compute-intensive prefill and memory-bound decode phases across specialized GPUs. This separation improves utilization and throughput under fixed hardware capacity. However, as model and cluster scales grow, power, rather than compute, has become the dominant limiter of overall performance and cost efficiency. In this paper, we propose RAPID, a power-aware disaggregated inference framework that jointly manages GPU roles and power budgets to sustain goodput within strict power caps. RAPID utilizes static and dynamic power reallocation in addition to GPU reallocation to improve performance under fixed power bounds. RAPID improves overall performance and application consistency beyond what is achievable in current disaggregation solutions, resulting in up to a 2x improvement in SLO attainment at peak load when compared to a static assignment without an increase in complexity or cost.
title Power Aware Dynamic Reallocation For Inference
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.12241