A Risk-Aware Reinforcement Learning Reward for Financial Trading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srivastava, Uditansh, Aryan, Shivam, Singh, Shaurya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915327093243904
author Srivastava, Uditansh
Aryan, Shivam
Singh, Shaurya
author_facet Srivastava, Uditansh
Aryan, Shivam
Singh, Shaurya
contents We propose a novel composite reward function for reinforcement learning in financial trading that balances return and risk using four differentiable terms: annualized return downside risk differential return and the Treynor ratio Unlike single metric objectives for example the Sharpe ratio our formulation is modular and parameterized by weights w1 w2 w3 and w4 enabling practitioners to encode diverse investor preferences We tune these weights via grid search to target specific risk return profiles We derive closed form gradients for each term to facilitate gradient based training and analyze key theoretical properties including monotonicity boundedness and modularity This framework offers a general blueprint for building robust multi objective reward functions in complex trading environments and can be extended with additional risk measures or adaptive weighting
format Preprint
id arxiv_https___arxiv_org_abs_2506_04358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Risk-Aware Reinforcement Learning Reward for Financial Trading
Srivastava, Uditansh
Aryan, Shivam
Singh, Shaurya
Machine Learning
We propose a novel composite reward function for reinforcement learning in financial trading that balances return and risk using four differentiable terms: annualized return downside risk differential return and the Treynor ratio Unlike single metric objectives for example the Sharpe ratio our formulation is modular and parameterized by weights w1 w2 w3 and w4 enabling practitioners to encode diverse investor preferences We tune these weights via grid search to target specific risk return profiles We derive closed form gradients for each term to facilitate gradient based training and analyze key theoretical properties including monotonicity boundedness and modularity This framework offers a general blueprint for building robust multi objective reward functions in complex trading environments and can be extended with additional risk measures or adaptive weighting
title A Risk-Aware Reinforcement Learning Reward for Financial Trading
topic Machine Learning
url https://arxiv.org/abs/2506.04358