Reinforcement Learning for Optimal Execution when Liquidity is Time-Varying

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Macrì, Andrea, Lillo, Fabrizio
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909112978112512
author Macrì, Andrea
Lillo, Fabrizio
author_facet Macrì, Andrea
Lillo, Fabrizio
contents Optimal execution is an important problem faced by any trader. Most solutions are based on the assumption of constant market impact, while liquidity is known to be dynamic. Moreover, models with time-varying liquidity typically assume that it is observable, despite the fact that, in reality, it is latent and hard to measure in real time. In this paper we show that the use of Double Deep Q-learning, a form of Reinforcement Learning based on neural networks, is able to learn optimal trading policies when liquidity is time-varying. Specifically, we consider an Almgren-Chriss framework with temporary and permanent impact parameters following several deterministic and stochastic dynamics. Using extensive numerical experiments, we show that the trained algorithm learns the optimal policy when the analytical solution is available, and overcomes benchmarks and approximated solutions when the solution is not available.
format Preprint
id arxiv_https___arxiv_org_abs_2402_12049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reinforcement Learning for Optimal Execution when Liquidity is Time-Varying
Macrì, Andrea
Lillo, Fabrizio
Trading and Market Microstructure
Optimal execution is an important problem faced by any trader. Most solutions are based on the assumption of constant market impact, while liquidity is known to be dynamic. Moreover, models with time-varying liquidity typically assume that it is observable, despite the fact that, in reality, it is latent and hard to measure in real time. In this paper we show that the use of Double Deep Q-learning, a form of Reinforcement Learning based on neural networks, is able to learn optimal trading policies when liquidity is time-varying. Specifically, we consider an Almgren-Chriss framework with temporary and permanent impact parameters following several deterministic and stochastic dynamics. Using extensive numerical experiments, we show that the trained algorithm learns the optimal policy when the analytical solution is available, and overcomes benchmarks and approximated solutions when the solution is not available.
title Reinforcement Learning for Optimal Execution when Liquidity is Time-Varying
topic Trading and Market Microstructure
url https://arxiv.org/abs/2402.12049