_version_ 1866916911317516288
author Choukse, Esha
Warrier, Brijesh
Heath, Scot
Belmont, Luz
Zhao, April
Khan, Hassan Ali
Harry, Brian
Kappel, Matthew
Hewett, Russell J.
Datta, Kushal
Pei, Yu
Lichtenberger, Caroline
Siegler, John
Lukofsky, David
Kahn, Zaid
Sahota, Gurpreet
Sullivan, Andy
Frederick, Charles
Thai, Hien
Naughton, Rebecca
Jurnove, Daniel
Harp, Justin
Carper, Reid
Mahalingam, Nithish
Varkala, Srini
Kumbhare, Alok Gautam
Desai, Satyajit
Ramamurthy, Venkatesh
Gottumukkala, Praneeth
Bhatia, Girish
Wildstone, Kelsey
Olariu, Laurentiu
Incorvaia, Ileana
Wetmore, Alex
Ram, Prabhat
Raghuraman, Melur
Ayna, Mohammed
Kendrick, Mike
Bianchini, Ricardo
Hurst, Aaron
Zamani, Reza
Li, Xin
Petrov, Michael
Oden, Gene
Carmichael, Rory
Li, Tom
Gupta, Apoorv
Patel, Pratikkumar
Dattani, Nilesh
Marwong, Lawrence
Nertney, Rob
Kobayashi, Hirofumi
Liott, Jeff
Enev, Miro
Ramakrishnan, Divya
Buck, Ian
Alben, Jonah
author_facet Choukse, Esha
Warrier, Brijesh
Heath, Scot
Belmont, Luz
Zhao, April
Khan, Hassan Ali
Harry, Brian
Kappel, Matthew
Hewett, Russell J.
Datta, Kushal
Pei, Yu
Lichtenberger, Caroline
Siegler, John
Lukofsky, David
Kahn, Zaid
Sahota, Gurpreet
Sullivan, Andy
Frederick, Charles
Thai, Hien
Naughton, Rebecca
Jurnove, Daniel
Harp, Justin
Carper, Reid
Mahalingam, Nithish
Varkala, Srini
Kumbhare, Alok Gautam
Desai, Satyajit
Ramamurthy, Venkatesh
Gottumukkala, Praneeth
Bhatia, Girish
Wildstone, Kelsey
Olariu, Laurentiu
Incorvaia, Ileana
Wetmore, Alex
Ram, Prabhat
Raghuraman, Melur
Ayna, Mohammed
Kendrick, Mike
Bianchini, Ricardo
Hurst, Aaron
Zamani, Reza
Li, Xin
Petrov, Michael
Oden, Gene
Carmichael, Rory
Li, Tom
Gupta, Apoorv
Patel, Pratikkumar
Dattani, Nilesh
Marwong, Lawrence
Nertney, Rob
Kobayashi, Hirofumi
Liott, Jeff
Enev, Miro
Ramakrishnan, Divya
Buck, Ian
Alben, Jonah
contents Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability in power consumption during the training. Given the synchronous nature of these jobs, during every iteration there is a computation-heavy phase, where each GPU works on the local data, and a communication-heavy phase where all the GPUs synchronize on the data. Because compute-heavy phases require much more power than communication phases, large power swings occur. The amplitude of these power swings is ever increasing with the increase in the size of training jobs. An even bigger challenge arises from the frequency spectrum of these power swings which, if harmonized with critical frequencies of utilities, can cause physical damage to the power grid infrastructure. Therefore, to continue scaling AI training workloads safely, we need to stabilize the power of such workloads. This paper introduces the challenge with production data and explores innovative solutions across the stack: software, GPU hardware, and datacenter infrastructure. We present the pros and cons of each of these approaches and finally present a multi-pronged approach to solving the challenge. The proposed solutions are rigorously tested using a combination of real hardware and Microsoft's in-house cloud power simulator, providing critical insights into the efficacy of these interventions under real-world conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Power Stabilization for AI Training Datacenters
Choukse, Esha
Warrier, Brijesh
Heath, Scot
Belmont, Luz
Zhao, April
Khan, Hassan Ali
Harry, Brian
Kappel, Matthew
Hewett, Russell J.
Datta, Kushal
Pei, Yu
Lichtenberger, Caroline
Siegler, John
Lukofsky, David
Kahn, Zaid
Sahota, Gurpreet
Sullivan, Andy
Frederick, Charles
Thai, Hien
Naughton, Rebecca
Jurnove, Daniel
Harp, Justin
Carper, Reid
Mahalingam, Nithish
Varkala, Srini
Kumbhare, Alok Gautam
Desai, Satyajit
Ramamurthy, Venkatesh
Gottumukkala, Praneeth
Bhatia, Girish
Wildstone, Kelsey
Olariu, Laurentiu
Incorvaia, Ileana
Wetmore, Alex
Ram, Prabhat
Raghuraman, Melur
Ayna, Mohammed
Kendrick, Mike
Bianchini, Ricardo
Hurst, Aaron
Zamani, Reza
Li, Xin
Petrov, Michael
Oden, Gene
Carmichael, Rory
Li, Tom
Gupta, Apoorv
Patel, Pratikkumar
Dattani, Nilesh
Marwong, Lawrence
Nertney, Rob
Kobayashi, Hirofumi
Liott, Jeff
Enev, Miro
Ramakrishnan, Divya
Buck, Ian
Alben, Jonah
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability in power consumption during the training. Given the synchronous nature of these jobs, during every iteration there is a computation-heavy phase, where each GPU works on the local data, and a communication-heavy phase where all the GPUs synchronize on the data. Because compute-heavy phases require much more power than communication phases, large power swings occur. The amplitude of these power swings is ever increasing with the increase in the size of training jobs. An even bigger challenge arises from the frequency spectrum of these power swings which, if harmonized with critical frequencies of utilities, can cause physical damage to the power grid infrastructure. Therefore, to continue scaling AI training workloads safely, we need to stabilize the power of such workloads. This paper introduces the challenge with production data and explores innovative solutions across the stack: software, GPU hardware, and datacenter infrastructure. We present the pros and cons of each of these approaches and finally present a multi-pronged approach to solving the challenge. The proposed solutions are rigorously tested using a combination of real hardware and Microsoft's in-house cloud power simulator, providing critical insights into the efficacy of these interventions under real-world conditions.
title Power Stabilization for AI Training Datacenters
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.14318