Time-Inhomogeneous Volatility Aversion for Financial Applications of Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cacciamani, Federico, Daluiso, Roberto, Pinciroli, Marco, Trapletti, Michele, Vittori, Edoardo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915794869288960
author Cacciamani, Federico
Daluiso, Roberto
Pinciroli, Marco
Trapletti, Michele
Vittori, Edoardo
author_facet Cacciamani, Federico
Daluiso, Roberto
Pinciroli, Marco
Trapletti, Michele
Vittori, Edoardo
contents In finance, sequential decision problems are often faced, for which reinforcement learning (RL) emerges as a promising tool for optimisation without the need of analytical tractability. However, the objective of classical RL is the expected cumulated reward, while financial applications typically require a trade-off between return and risk. In this work, we focus on settings where one cares about the time split of the total return, ruling out most risk-aware generalisations of RL which optimise a risk measure defined on the latter. We notice that a preference for homogeneous splits, which we found satisfactory for hedging, can be unfit for other problems, and therefore propose a new risk metric which still penalises uncertainty of the single rewards, but allows for an arbitrary planning of their target levels. We study the properties of the resulting objective and the generalisation of learning algorithms to optimise it. Finally, we show numerical results on toy examples.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12030
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Time-Inhomogeneous Volatility Aversion for Financial Applications of Reinforcement Learning
Cacciamani, Federico
Daluiso, Roberto
Pinciroli, Marco
Trapletti, Michele
Vittori, Edoardo
Computational Finance
Trading and Market Microstructure
91-08
In finance, sequential decision problems are often faced, for which reinforcement learning (RL) emerges as a promising tool for optimisation without the need of analytical tractability. However, the objective of classical RL is the expected cumulated reward, while financial applications typically require a trade-off between return and risk. In this work, we focus on settings where one cares about the time split of the total return, ruling out most risk-aware generalisations of RL which optimise a risk measure defined on the latter. We notice that a preference for homogeneous splits, which we found satisfactory for hedging, can be unfit for other problems, and therefore propose a new risk metric which still penalises uncertainty of the single rewards, but allows for an arbitrary planning of their target levels. We study the properties of the resulting objective and the generalisation of learning algorithms to optimise it. Finally, we show numerical results on toy examples.
title Time-Inhomogeneous Volatility Aversion for Financial Applications of Reinforcement Learning
topic Computational Finance
Trading and Market Microstructure
91-08
url https://arxiv.org/abs/2602.12030