Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Espana, Tomas, Hafsi, Yadh, Lillo, Fabrizio, Vittori, Edoardo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911276415844352
author Espana, Tomas
Hafsi, Yadh
Lillo, Fabrizio
Vittori, Edoardo
author_facet Espana, Tomas
Hafsi, Yadh
Lillo, Fabrizio
Vittori, Edoardo
contents We investigate the use of Reinforcement Learning for the optimal execution of meta-orders, where the objective is to execute incrementally large orders while minimizing implementation shortfall and market impact over an extended period of time. Departing from traditional parametric approaches to price dynamics and impact modeling, we adopt a model-free, data-driven framework. Since policy optimization requires counterfactual feedback that historical data cannot provide, we employ the Queue-Reactive Model to generate realistic and tractable limit order book simulations that encompass transient price impact, and nonlinear and dynamic order flow responses. Methodologically, we train a Double Deep Q-Network agent on a state space comprising time, inventory, price, and depth variables, and evaluate its performance against established benchmarks. Numerical simulation results show that the agent learns a policy that is both strategic and tactical, adapting effectively to order book conditions and outperforming standard approaches across multiple training configurations. These findings provide strong evidence that model-free Reinforcement Learning can yield adaptive and robust solutions to the optimal execution problem.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution
Espana, Tomas
Hafsi, Yadh
Lillo, Fabrizio
Vittori, Edoardo
Trading and Market Microstructure
Machine Learning
We investigate the use of Reinforcement Learning for the optimal execution of meta-orders, where the objective is to execute incrementally large orders while minimizing implementation shortfall and market impact over an extended period of time. Departing from traditional parametric approaches to price dynamics and impact modeling, we adopt a model-free, data-driven framework. Since policy optimization requires counterfactual feedback that historical data cannot provide, we employ the Queue-Reactive Model to generate realistic and tractable limit order book simulations that encompass transient price impact, and nonlinear and dynamic order flow responses. Methodologically, we train a Double Deep Q-Network agent on a state space comprising time, inventory, price, and depth variables, and evaluate its performance against established benchmarks. Numerical simulation results show that the agent learns a policy that is both strategic and tactical, adapting effectively to order book conditions and outperforming standard approaches across multiple training configurations. These findings provide strong evidence that model-free Reinforcement Learning can yield adaptive and robust solutions to the optimal execution problem.
title Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution
topic Trading and Market Microstructure
Machine Learning
url https://arxiv.org/abs/2511.15262