Power Stabilization for AI Training Datacenters
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916911317516288 |
|---|---|
| author | Choukse, Esha Warrier, Brijesh Heath, Scot Belmont, Luz Zhao, April Khan, Hassan Ali Harry, Brian Kappel, Matthew Hewett, Russell J. Datta, Kushal Pei, Yu Lichtenberger, Caroline Siegler, John Lukofsky, David Kahn, Zaid Sahota, Gurpreet Sullivan, Andy Frederick, Charles Thai, Hien Naughton, Rebecca Jurnove, Daniel Harp, Justin Carper, Reid Mahalingam, Nithish Varkala, Srini Kumbhare, Alok Gautam Desai, Satyajit Ramamurthy, Venkatesh Gottumukkala, Praneeth Bhatia, Girish Wildstone, Kelsey Olariu, Laurentiu Incorvaia, Ileana Wetmore, Alex Ram, Prabhat Raghuraman, Melur Ayna, Mohammed Kendrick, Mike Bianchini, Ricardo Hurst, Aaron Zamani, Reza Li, Xin Petrov, Michael Oden, Gene Carmichael, Rory Li, Tom Gupta, Apoorv Patel, Pratikkumar Dattani, Nilesh Marwong, Lawrence Nertney, Rob Kobayashi, Hirofumi Liott, Jeff Enev, Miro Ramakrishnan, Divya Buck, Ian Alben, Jonah |
| author_facet | Choukse, Esha Warrier, Brijesh Heath, Scot Belmont, Luz Zhao, April Khan, Hassan Ali Harry, Brian Kappel, Matthew Hewett, Russell J. Datta, Kushal Pei, Yu Lichtenberger, Caroline Siegler, John Lukofsky, David Kahn, Zaid Sahota, Gurpreet Sullivan, Andy Frederick, Charles Thai, Hien Naughton, Rebecca Jurnove, Daniel Harp, Justin Carper, Reid Mahalingam, Nithish Varkala, Srini Kumbhare, Alok Gautam Desai, Satyajit Ramamurthy, Venkatesh Gottumukkala, Praneeth Bhatia, Girish Wildstone, Kelsey Olariu, Laurentiu Incorvaia, Ileana Wetmore, Alex Ram, Prabhat Raghuraman, Melur Ayna, Mohammed Kendrick, Mike Bianchini, Ricardo Hurst, Aaron Zamani, Reza Li, Xin Petrov, Michael Oden, Gene Carmichael, Rory Li, Tom Gupta, Apoorv Patel, Pratikkumar Dattani, Nilesh Marwong, Lawrence Nertney, Rob Kobayashi, Hirofumi Liott, Jeff Enev, Miro Ramakrishnan, Divya Buck, Ian Alben, Jonah |
| contents | Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability in power consumption during the training. Given the synchronous nature of these jobs, during every iteration there is a computation-heavy phase, where each GPU works on the local data, and a communication-heavy phase where all the GPUs synchronize on the data. Because compute-heavy phases require much more power than communication phases, large power swings occur. The amplitude of these power swings is ever increasing with the increase in the size of training jobs. An even bigger challenge arises from the frequency spectrum of these power swings which, if harmonized with critical frequencies of utilities, can cause physical damage to the power grid infrastructure. Therefore, to continue scaling AI training workloads safely, we need to stabilize the power of such workloads. This paper introduces the challenge with production data and explores innovative solutions across the stack: software, GPU hardware, and datacenter infrastructure. We present the pros and cons of each of these approaches and finally present a multi-pronged approach to solving the challenge. The proposed solutions are rigorously tested using a combination of real hardware and Microsoft's in-house cloud power simulator, providing critical insights into the efficacy of these interventions under real-world conditions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_14318 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Power Stabilization for AI Training Datacenters Choukse, Esha Warrier, Brijesh Heath, Scot Belmont, Luz Zhao, April Khan, Hassan Ali Harry, Brian Kappel, Matthew Hewett, Russell J. Datta, Kushal Pei, Yu Lichtenberger, Caroline Siegler, John Lukofsky, David Kahn, Zaid Sahota, Gurpreet Sullivan, Andy Frederick, Charles Thai, Hien Naughton, Rebecca Jurnove, Daniel Harp, Justin Carper, Reid Mahalingam, Nithish Varkala, Srini Kumbhare, Alok Gautam Desai, Satyajit Ramamurthy, Venkatesh Gottumukkala, Praneeth Bhatia, Girish Wildstone, Kelsey Olariu, Laurentiu Incorvaia, Ileana Wetmore, Alex Ram, Prabhat Raghuraman, Melur Ayna, Mohammed Kendrick, Mike Bianchini, Ricardo Hurst, Aaron Zamani, Reza Li, Xin Petrov, Michael Oden, Gene Carmichael, Rory Li, Tom Gupta, Apoorv Patel, Pratikkumar Dattani, Nilesh Marwong, Lawrence Nertney, Rob Kobayashi, Hirofumi Liott, Jeff Enev, Miro Ramakrishnan, Divya Buck, Ian Alben, Jonah Hardware Architecture Artificial Intelligence Distributed, Parallel, and Cluster Computing Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability in power consumption during the training. Given the synchronous nature of these jobs, during every iteration there is a computation-heavy phase, where each GPU works on the local data, and a communication-heavy phase where all the GPUs synchronize on the data. Because compute-heavy phases require much more power than communication phases, large power swings occur. The amplitude of these power swings is ever increasing with the increase in the size of training jobs. An even bigger challenge arises from the frequency spectrum of these power swings which, if harmonized with critical frequencies of utilities, can cause physical damage to the power grid infrastructure. Therefore, to continue scaling AI training workloads safely, we need to stabilize the power of such workloads. This paper introduces the challenge with production data and explores innovative solutions across the stack: software, GPU hardware, and datacenter infrastructure. We present the pros and cons of each of these approaches and finally present a multi-pronged approach to solving the challenge. The proposed solutions are rigorously tested using a combination of real hardware and Microsoft's in-house cloud power simulator, providing critical insights into the efficacy of these interventions under real-world conditions. |
| title | Power Stabilization for AI Training Datacenters |
| topic | Hardware Architecture Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2508.14318 |