Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Majumder, Subir, Yu, Minlan, Xie, Le
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908957616898048
author Majumder, Subir
Yu, Minlan
Xie, Le
author_facet Majumder, Subir
Yu, Minlan
Xie, Le
contents Artificial intelligence (AI) is driving rapid growth in electricity demand, yet the grid-facing power dynamics of AI data centers remain poorly understood. Here we show that, in shared-GPU systems, the composition of batch and inference workloads decouples aggregate power variability from short-horizon ramping. As the inference share rises, variability becomes U-shaped, whereas ramping becomes hump-shaped, particularly under higher loading. The magnitude and turning points of these patterns also depend on system loading. Using a trace-calibrated framework linking workload arrivals, queueing, scheduling, and GPU power, we show that the underlying mechanism is asymmetric. At intermediate workload mixes, queued batch jobs fill capacity left idle by fluctuating inference demand, reducing aggregate power variability. However, short-horizon ramping remains elevated because inference-side fluctuations propagate more directly into realized power. AI data centers should therefore be understood as dynamic systems whose workload composition shapes their grid impact.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10769
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers
Majumder, Subir
Yu, Minlan
Xie, Le
Systems and Control
Distributed, Parallel, and Cluster Computing
Performance
Artificial intelligence (AI) is driving rapid growth in electricity demand, yet the grid-facing power dynamics of AI data centers remain poorly understood. Here we show that, in shared-GPU systems, the composition of batch and inference workloads decouples aggregate power variability from short-horizon ramping. As the inference share rises, variability becomes U-shaped, whereas ramping becomes hump-shaped, particularly under higher loading. The magnitude and turning points of these patterns also depend on system loading. Using a trace-calibrated framework linking workload arrivals, queueing, scheduling, and GPU power, we show that the underlying mechanism is asymmetric. At intermediate workload mixes, queued batch jobs fill capacity left idle by fluctuating inference demand, reducing aggregate power variability. However, short-horizon ramping remains elevated because inference-side fluctuations propagate more directly into realized power. AI data centers should therefore be understood as dynamic systems whose workload composition shapes their grid impact.
title Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers
topic Systems and Control
Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2604.10769