A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917410861219840 |
|---|---|
| author | Razumovskiy, Lev Karenin, Nikolay |
| author_facet | Razumovskiy, Lev Karenin, Nikolay |
| contents | This paper provides a systematic comparison between Fitted Dynamic Programming (DP), where demand is estimated from data, and Reinforcement Learning (RL) methods in finite-horizon dynamic pricing problems. We analyze their performance across environments of increasing structural complexity, ranging from a single typology benchmark to multi-typology settings with heterogeneous demand and inter-temporal revenue constraints. Unlike simplified comparisons that restrict DP to low-dimensional settings, we apply dynamic programming in richer, multi-dimensional environments with multiple product types and constraints. We evaluate revenue performance, stability, constraint satisfaction behavior, and computational scaling, highlighting the trade-offs between explicit expectation-based optimization and trajectory-based learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_14059 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing Razumovskiy, Lev Karenin, Nikolay General Economics Economics Machine Learning This paper provides a systematic comparison between Fitted Dynamic Programming (DP), where demand is estimated from data, and Reinforcement Learning (RL) methods in finite-horizon dynamic pricing problems. We analyze their performance across environments of increasing structural complexity, ranging from a single typology benchmark to multi-typology settings with heterogeneous demand and inter-temporal revenue constraints. Unlike simplified comparisons that restrict DP to low-dimensional settings, we apply dynamic programming in richer, multi-dimensional environments with multiple product types and constraints. We evaluate revenue performance, stability, constraint satisfaction behavior, and computational scaling, highlighting the trade-offs between explicit expectation-based optimization and trajectory-based learning. |
| title | A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing |
| topic | General Economics Economics Machine Learning |
| url | https://arxiv.org/abs/2604.14059 |