Myopic Optimality: why reinforcement learning portfolio management strategies lose money
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911157275590656 |
|---|---|
| author | Ma, Yuming |
| author_facet | Ma, Yuming |
| contents | Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dual gap and convergence results: geometric MO vs. RL floors. We quantify phantom profit in RL via Malliavin policy-gradient contamination analysis and define a control-affects-dynamics (CAD) premium of RL indicating plausibly positive. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_12764 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Myopic Optimality: why reinforcement learning portfolio management strategies lose money Ma, Yuming Trading and Market Microstructure Optimization and Control Probability Portfolio Management Risk Management 91G10, 60H07, 90C25, 91G80, 93E20, 60H10 Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dual gap and convergence results: geometric MO vs. RL floors. We quantify phantom profit in RL via Malliavin policy-gradient contamination analysis and define a control-affects-dynamics (CAD) premium of RL indicating plausibly positive. |
| title | Myopic Optimality: why reinforcement learning portfolio management strategies lose money |
| topic | Trading and Market Microstructure Optimization and Control Probability Portfolio Management Risk Management 91G10, 60H07, 90C25, 91G80, 93E20, 60H10 |
| url | https://arxiv.org/abs/2509.12764 |