Myopic Optimality: why reinforcement learning portfolio management strategies lose money

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Ma, Yuming
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911157275590656
author Ma, Yuming
author_facet Ma, Yuming
contents Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dual gap and convergence results: geometric MO vs. RL floors. We quantify phantom profit in RL via Malliavin policy-gradient contamination analysis and define a control-affects-dynamics (CAD) premium of RL indicating plausibly positive.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Myopic Optimality: why reinforcement learning portfolio management strategies lose money
Ma, Yuming
Trading and Market Microstructure
Optimization and Control
Probability
Portfolio Management
Risk Management
91G10, 60H07, 90C25, 91G80, 93E20, 60H10
Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dual gap and convergence results: geometric MO vs. RL floors. We quantify phantom profit in RL via Malliavin policy-gradient contamination analysis and define a control-affects-dynamics (CAD) premium of RL indicating plausibly positive.
title Myopic Optimality: why reinforcement learning portfolio management strategies lose money
topic Trading and Market Microstructure
Optimization and Control
Probability
Portfolio Management
Risk Management
91G10, 60H07, 90C25, 91G80, 93E20, 60H10
url https://arxiv.org/abs/2509.12764