A General Approach of Automated Environment Design for Learning the Optimal Power Flow

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wolgast, Thomas, Nieße, Astrid
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910940637691904
author Wolgast, Thomas
Nieße, Astrid
author_facet Wolgast, Thomas
Nieße, Astrid
contents Reinforcement learning (RL) algorithms are increasingly used to solve the optimal power flow (OPF) problem. Yet, the question of how to design RL environments to maximize training performance remains unanswered, both for the OPF and the general case. We propose a general approach for automated RL environment design by utilizing multi-objective optimization. For that, we use the hyperparameter optimization (HPO) framework, which allows the reuse of existing HPO algorithms and methods. On five OPF benchmark problems, we demonstrate that our automated design approach consistently outperforms a manually created baseline environment design. Further, we use statistical analyses to determine which environment design decisions are especially important for performance, resulting in multiple novel insights on how RL-OPF environments should be designed. Finally, we discuss the risk of overfitting the environment to the utilized RL algorithm. To the best of our knowledge, this is the first general approach for automated RL environment design.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A General Approach of Automated Environment Design for Learning the Optimal Power Flow
Wolgast, Thomas
Nieße, Astrid
Machine Learning
Artificial Intelligence
Systems and Control
Reinforcement learning (RL) algorithms are increasingly used to solve the optimal power flow (OPF) problem. Yet, the question of how to design RL environments to maximize training performance remains unanswered, both for the OPF and the general case. We propose a general approach for automated RL environment design by utilizing multi-objective optimization. For that, we use the hyperparameter optimization (HPO) framework, which allows the reuse of existing HPO algorithms and methods. On five OPF benchmark problems, we demonstrate that our automated design approach consistently outperforms a manually created baseline environment design. Further, we use statistical analyses to determine which environment design decisions are especially important for performance, resulting in multiple novel insights on how RL-OPF environments should be designed. Finally, we discuss the risk of overfitting the environment to the utilized RL algorithm. To the best of our knowledge, this is the first general approach for automated RL environment design.
title A General Approach of Automated Environment Design for Learning the Optimal Power Flow
topic Machine Learning
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2505.07832