The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiashun, Obando-Ceron, Johan, Castro, Pablo Samuel, Courville, Aaron, Pan, Ling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912432881926144
author Liu, Jiashun
Obando-Ceron, Johan
Castro, Pablo Samuel
Courville, Aaron
Pan, Ling
author_facet Liu, Jiashun
Obando-Ceron, Johan
Castro, Pablo Samuel
Courville, Aaron
Pan, Ling
contents Off-policy deep reinforcement learning (RL) typically leverages replay buffers for reusing past experiences during learning. This can help improve sample efficiency when the collected data is informative and aligned with the learning objectives; when that is not the case, it can have the effect of "polluting" the replay buffer with data which can exacerbate optimization challenges in addition to wasting environment interactions due to wasteful sampling. We argue that sampling these uninformative and wasteful transitions can be avoided by addressing the sunk cost fallacy, which, in the context of deep RL, is the tendency towards continuing an episode until termination. To address this, we propose learn to stop (LEAST), a lightweight mechanism that enables strategic early episode termination based on Q-value and gradient statistics, which helps agents recognize when to terminate unproductive episodes early. We demonstrate that our method improves learning efficiency on a variety of RL algorithms, evaluated on both the MuJoCo and DeepMind Control Suite benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
Liu, Jiashun
Obando-Ceron, Johan
Castro, Pablo Samuel
Courville, Aaron
Pan, Ling
Machine Learning
Off-policy deep reinforcement learning (RL) typically leverages replay buffers for reusing past experiences during learning. This can help improve sample efficiency when the collected data is informative and aligned with the learning objectives; when that is not the case, it can have the effect of "polluting" the replay buffer with data which can exacerbate optimization challenges in addition to wasting environment interactions due to wasteful sampling. We argue that sampling these uninformative and wasteful transitions can be avoided by addressing the sunk cost fallacy, which, in the context of deep RL, is the tendency towards continuing an episode until termination. To address this, we propose learn to stop (LEAST), a lightweight mechanism that enables strategic early episode termination based on Q-value and gradient statistics, which helps agents recognize when to terminate unproductive episodes early. We demonstrate that our method improves learning efficiency on a variety of RL algorithms, evaluated on both the MuJoCo and DeepMind Control Suite benchmarks.
title The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2506.13672