Learning Reachability of Energy Storage Arbitrage

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tapia, Tomás, Castellano, Agustin, Mallada, Enrique, Dvorkin, Yury
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918489503039488
author Tapia, Tomás
Castellano, Agustin
Mallada, Enrique
Dvorkin, Yury
author_facet Tapia, Tomás
Castellano, Agustin
Mallada, Enrique
Dvorkin, Yury
contents Power systems face increasing weather-driven variability and, therefore, increasingly rely on flexible but energy-limited storage resources. Energy storage can buffer this variability, but its value depends on intertemporal decisions under uncertain prices. Without accounting for the future reliability value of stored energy, batteries may act myopically, discharging too early or failing to preserve reserves during critical hours. This paper introduces a stopping-time reward that, together with a state-of-charge (SoC) range target penalty, aligns arbitrage incentives with system reliability by rewarding storage that maintains sufficient SoC before critical hours. We formulate the problem as an online optimization with a chance-constrained terminal SoC and embed it in an end-to-end (E2E) learning framework, jointly training the price predictor and control policy. The proposed design enhances reachability of target SoC ranges, improves profit under volatile conditions, and reduces its standard deviation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Reachability of Energy Storage Arbitrage
Tapia, Tomás
Castellano, Agustin
Mallada, Enrique
Dvorkin, Yury
Systems and Control
Optimization and Control
Power systems face increasing weather-driven variability and, therefore, increasingly rely on flexible but energy-limited storage resources. Energy storage can buffer this variability, but its value depends on intertemporal decisions under uncertain prices. Without accounting for the future reliability value of stored energy, batteries may act myopically, discharging too early or failing to preserve reserves during critical hours. This paper introduces a stopping-time reward that, together with a state-of-charge (SoC) range target penalty, aligns arbitrage incentives with system reliability by rewarding storage that maintains sufficient SoC before critical hours. We formulate the problem as an online optimization with a chance-constrained terminal SoC and embed it in an end-to-end (E2E) learning framework, jointly training the price predictor and control policy. The proposed design enhances reachability of target SoC ranges, improves profit under volatile conditions, and reduces its standard deviation.
title Learning Reachability of Energy Storage Arbitrage
topic Systems and Control
Optimization and Control
url https://arxiv.org/abs/2512.06600