An approach of deep reinforcement learning for maximizing the net present value of stochastic projects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Wei, Yang, Fan, Cui, Qinyuan, Chen, Zhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917084840067072
author Xu, Wei
Yang, Fan
Cui, Qinyuan
Chen, Zhi
author_facet Xu, Wei
Yang, Fan
Cui, Qinyuan
Chen, Zhi
contents This paper investigates a project with stochastic activity durations and cash flows under discrete scenarios, where activities must satisfy precedence constraints generating cash inflows and outflows. The objective is to maximize expected net present value (NPV) by accelerating inflows and deferring outflows. We formulate the problem as a discrete-time Markov Decision Process (MDP) and propose a Double Deep Q-Network (DDQN) approach. Comparative experiments demonstrate that DDQN outperforms traditional rigid and dynamic strategies, particularly in large-scale or highly uncertain environments, exhibiting superior computational capability, policy reliability, and adaptability. Ablation studies further reveal that the dual-network architecture mitigates overestimation of action values, while the target network substantially improves training convergence and robustness. These results indicate that DDQN not only achieves higher expected NPV in complex project optimization but also provides a reliable framework for stable and effective policy implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12865
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
Xu, Wei
Yang, Fan
Cui, Qinyuan
Chen, Zhi
Machine Learning
Artificial Intelligence
This paper investigates a project with stochastic activity durations and cash flows under discrete scenarios, where activities must satisfy precedence constraints generating cash inflows and outflows. The objective is to maximize expected net present value (NPV) by accelerating inflows and deferring outflows. We formulate the problem as a discrete-time Markov Decision Process (MDP) and propose a Double Deep Q-Network (DDQN) approach. Comparative experiments demonstrate that DDQN outperforms traditional rigid and dynamic strategies, particularly in large-scale or highly uncertain environments, exhibiting superior computational capability, policy reliability, and adaptability. Ablation studies further reveal that the dual-network architecture mitigates overestimation of action values, while the target network substantially improves training convergence and robustness. These results indicate that DDQN not only achieves higher expected NPV in complex project optimization but also provides a reliable framework for stable and effective policy implementation.
title An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.12865