Risk-Aware Deep Reinforcement Learning for Dynamic Portfolio Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lwele, Emmanuel, Emmanuel, Sabuni, Sitali, Sitali Gabriel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911266119876608
author Lwele, Emmanuel
Emmanuel, Sabuni
Sitali, Sitali Gabriel
author_facet Lwele, Emmanuel
Emmanuel, Sabuni
Sitali, Sitali Gabriel
contents This paper presents a deep reinforcement learning (DRL) framework for dynamic portfolio optimization under market uncertainty and risk. The proposed model integrates a Sharpe ratio-based reward function with direct risk control mechanisms, including maximum drawdown and volatility constraints. Proximal Policy Optimization (PPO) is employed to learn adaptive asset allocation strategies over historical financial time series. Model performance is benchmarked against mean-variance and equal-weight portfolio strategies using backtesting on high-performing equities. Results indicate that the DRL agent stabilizes volatility successfully but suffers from degraded risk-adjusted returns due to over-conservative policy convergence, highlighting the challenge of balancing exploration, return maximization, and risk mitigation. The study underscores the need for improved reward shaping and hybrid risk-aware strategies to enhance the practical deployment of DRL-based portfolio allocation models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11481
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Risk-Aware Deep Reinforcement Learning for Dynamic Portfolio Optimization
Lwele, Emmanuel
Emmanuel, Sabuni
Sitali, Sitali Gabriel
Portfolio Management
Computational Engineering, Finance, and Science
Econometrics
91G10, 91B84, 68T07
I.2.6; I.2.7; J.4
This paper presents a deep reinforcement learning (DRL) framework for dynamic portfolio optimization under market uncertainty and risk. The proposed model integrates a Sharpe ratio-based reward function with direct risk control mechanisms, including maximum drawdown and volatility constraints. Proximal Policy Optimization (PPO) is employed to learn adaptive asset allocation strategies over historical financial time series. Model performance is benchmarked against mean-variance and equal-weight portfolio strategies using backtesting on high-performing equities. Results indicate that the DRL agent stabilizes volatility successfully but suffers from degraded risk-adjusted returns due to over-conservative policy convergence, highlighting the challenge of balancing exploration, return maximization, and risk mitigation. The study underscores the need for improved reward shaping and hybrid risk-aware strategies to enhance the practical deployment of DRL-based portfolio allocation models.
title Risk-Aware Deep Reinforcement Learning for Dynamic Portfolio Optimization
topic Portfolio Management
Computational Engineering, Finance, and Science
Econometrics
91G10, 91B84, 68T07
I.2.6; I.2.7; J.4
url https://arxiv.org/abs/2511.11481