Economic Battery Storage Dispatch with Deep Reinforcement Learning from Rule-Based Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sage, Manuel, Staniszewski, Martin, Zhao, Yaoyao Fiona
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917978682949632
author Sage, Manuel
Staniszewski, Martin
Zhao, Yaoyao Fiona
author_facet Sage, Manuel
Staniszewski, Martin
Zhao, Yaoyao Fiona
contents The application of deep reinforcement learning algorithms to economic battery dispatch problems has significantly increased recently. However, optimizing battery dispatch over long horizons can be challenging due to delayed rewards. In our experiments we observe poor performance of popular actor-critic algorithms when trained on yearly episodes with hourly resolution. To address this, we propose an approach extending soft actor-critic (SAC) with learning from demonstrations. The special feature of our approach is that, due to the absence of expert demonstrations, the demonstration data is generated through simple, rule-based policies. We conduct a case study on a grid-connected microgrid and use if-then-else statements based on the wholesale price of electricity to collect demonstrations. These are stored in a separate replay buffer and sampled with linearly decaying probability along with the agent's own experiences. Despite these minimal modifications and the imperfections in the demonstration data, the results show a drastic performance improvement regarding both sample efficiency and final rewards. We further show that the proposed method reliably outperforms the demonstrator and is robust to the choice of rule, as long as the rule is sufficient to guide early training into the right direction.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04326
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Economic Battery Storage Dispatch with Deep Reinforcement Learning from Rule-Based Demonstrations
Sage, Manuel
Staniszewski, Martin
Zhao, Yaoyao Fiona
Systems and Control
Machine Learning
The application of deep reinforcement learning algorithms to economic battery dispatch problems has significantly increased recently. However, optimizing battery dispatch over long horizons can be challenging due to delayed rewards. In our experiments we observe poor performance of popular actor-critic algorithms when trained on yearly episodes with hourly resolution. To address this, we propose an approach extending soft actor-critic (SAC) with learning from demonstrations. The special feature of our approach is that, due to the absence of expert demonstrations, the demonstration data is generated through simple, rule-based policies. We conduct a case study on a grid-connected microgrid and use if-then-else statements based on the wholesale price of electricity to collect demonstrations. These are stored in a separate replay buffer and sampled with linearly decaying probability along with the agent's own experiences. Despite these minimal modifications and the imperfections in the demonstration data, the results show a drastic performance improvement regarding both sample efficiency and final rewards. We further show that the proposed method reliably outperforms the demonstrator and is robust to the choice of rule, as long as the rule is sufficient to guide early training into the right direction.
title Economic Battery Storage Dispatch with Deep Reinforcement Learning from Rule-Based Demonstrations
topic Systems and Control
Machine Learning
url https://arxiv.org/abs/2504.04326