Robot See, Robot Do: Imitation Reward for Noisy Financial Environments

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Goluža, Sven, Kovačević, Tomislav, Begušić, Stjepan, Kostanjčar, Zvonko
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917836028379136
author Goluža, Sven
Kovačević, Tomislav
Begušić, Stjepan
Kostanjčar, Zvonko
author_facet Goluža, Sven
Kovačević, Tomislav
Begušić, Stjepan
Kostanjčar, Zvonko
contents The sequential nature of decision-making in financial asset trading aligns naturally with the reinforcement learning (RL) framework, making RL a common approach in this domain. However, the low signal-to-noise ratio in financial markets results in noisy estimates of environment components, including the reward function, which hinders effective policy learning by RL agents. Given the critical importance of reward function design in RL problems, this paper introduces a novel and more robust reward function by leveraging imitation learning, where a trend labeling algorithm acts as an expert. We integrate imitation (expert's) feedback with reinforcement (agent's) feedback in a model-free RL algorithm, effectively embedding the imitation learning problem within the RL paradigm to handle the stochasticity of reward signals. Empirical results demonstrate that this novel approach improves financial performance metrics compared to traditional benchmarks and RL agents trained solely using reinforcement feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08637
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robot See, Robot Do: Imitation Reward for Noisy Financial Environments
Goluža, Sven
Kovačević, Tomislav
Begušić, Stjepan
Kostanjčar, Zvonko
Machine Learning
Robotics
Trading and Market Microstructure
The sequential nature of decision-making in financial asset trading aligns naturally with the reinforcement learning (RL) framework, making RL a common approach in this domain. However, the low signal-to-noise ratio in financial markets results in noisy estimates of environment components, including the reward function, which hinders effective policy learning by RL agents. Given the critical importance of reward function design in RL problems, this paper introduces a novel and more robust reward function by leveraging imitation learning, where a trend labeling algorithm acts as an expert. We integrate imitation (expert's) feedback with reinforcement (agent's) feedback in a model-free RL algorithm, effectively embedding the imitation learning problem within the RL paradigm to handle the stochasticity of reward signals. Empirical results demonstrate that this novel approach improves financial performance metrics compared to traditional benchmarks and RL agents trained solely using reinforcement feedback.
title Robot See, Robot Do: Imitation Reward for Noisy Financial Environments
topic Machine Learning
Robotics
Trading and Market Microstructure
url https://arxiv.org/abs/2411.08637