LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naug, Avisek, Guillen, Antonio, Kumar, Vineet, Greenwood, Scott, Brewer, Wesley, Ghorbanpour, Sahand, Babu, Ashwin Ramesh, Gundecha, Vineet, Gutierrez, Ricardo Luna, Sarkar, Soumyendu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914127510765568
author Naug, Avisek
Guillen, Antonio
Kumar, Vineet
Greenwood, Scott
Brewer, Wesley
Ghorbanpour, Sahand
Babu, Ashwin Ramesh
Gundecha, Vineet
Gutierrez, Ricardo Luna
Sarkar, Soumyendu
author_facet Naug, Avisek
Guillen, Antonio
Kumar, Vineet
Greenwood, Scott
Brewer, Wesley
Ghorbanpour, Sahand
Babu, Ashwin Ramesh
Gundecha, Vineet
Gutierrez, Ricardo Luna
Sarkar, Soumyendu
contents Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers
Naug, Avisek
Guillen, Antonio
Kumar, Vineet
Greenwood, Scott
Brewer, Wesley
Ghorbanpour, Sahand
Babu, Ashwin Ramesh
Gundecha, Vineet
Gutierrez, Ricardo Luna
Sarkar, Soumyendu
Machine Learning
Artificial Intelligence
Multiagent Systems
Systems and Control
Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.
title LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers
topic Machine Learning
Artificial Intelligence
Multiagent Systems
Systems and Control
url https://arxiv.org/abs/2511.00116