Time-Inhomogeneous Volatility Aversion for Financial Applications of Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915794869288960 |
|---|---|
| author | Cacciamani, Federico Daluiso, Roberto Pinciroli, Marco Trapletti, Michele Vittori, Edoardo |
| author_facet | Cacciamani, Federico Daluiso, Roberto Pinciroli, Marco Trapletti, Michele Vittori, Edoardo |
| contents | In finance, sequential decision problems are often faced, for which reinforcement learning (RL) emerges as a promising tool for optimisation without the need of analytical tractability. However, the objective of classical RL is the expected cumulated reward, while financial applications typically require a trade-off between return and risk. In this work, we focus on settings where one cares about the time split of the total return, ruling out most risk-aware generalisations of RL which optimise a risk measure defined on the latter. We notice that a preference for homogeneous splits, which we found satisfactory for hedging, can be unfit for other problems, and therefore propose a new risk metric which still penalises uncertainty of the single rewards, but allows for an arbitrary planning of their target levels. We study the properties of the resulting objective and the generalisation of learning algorithms to optimise it. Finally, we show numerical results on toy examples. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_12030 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Time-Inhomogeneous Volatility Aversion for Financial Applications of Reinforcement Learning Cacciamani, Federico Daluiso, Roberto Pinciroli, Marco Trapletti, Michele Vittori, Edoardo Computational Finance Trading and Market Microstructure 91-08 In finance, sequential decision problems are often faced, for which reinforcement learning (RL) emerges as a promising tool for optimisation without the need of analytical tractability. However, the objective of classical RL is the expected cumulated reward, while financial applications typically require a trade-off between return and risk. In this work, we focus on settings where one cares about the time split of the total return, ruling out most risk-aware generalisations of RL which optimise a risk measure defined on the latter. We notice that a preference for homogeneous splits, which we found satisfactory for hedging, can be unfit for other problems, and therefore propose a new risk metric which still penalises uncertainty of the single rewards, but allows for an arbitrary planning of their target levels. We study the properties of the resulting objective and the generalisation of learning algorithms to optimise it. Finally, we show numerical results on toy examples. |
| title | Time-Inhomogeneous Volatility Aversion for Financial Applications of Reinforcement Learning |
| topic | Computational Finance Trading and Market Microstructure 91-08 |
| url | https://arxiv.org/abs/2602.12030 |