SVL: Goal-Conditioned Reinforcement Learning as Survival Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tiofack, Franki Nguimatsia, Schramm, Fabian, Hellard, Théotime Le, Carpentier, Justin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910271213142016
author Tiofack, Franki Nguimatsia
Schramm, Fabian
Hellard, Théotime Le
Carpentier, Justin
author_facet Tiofack, Franki Nguimatsia
Schramm, Fabian
Hellard, Théotime Le
Carpentier, Justin
contents Standard approaches to goal-conditioned reinforcement learning (GCRL) that rely on temporal-difference learning can be unstable and sample-inefficient due to bootstrapping. While recent work has explored contrastive and supervised formulations to improve stability, we present a probabilistic alternative, called survival value learning (SVL), that reframes GCRL as a survival learning problem by modeling the time-to-goal from each state as a probability distribution. This structured distributional Monte Carlo perspective yields a closed-form identity that expresses the goal-conditioned value function as a discounted sum of survival probabilities, enabling value estimation via a hazard model trained via maximum likelihood on both event and right-censored trajectories. We introduce three practical value estimators, including finite-horizon truncation and two binned infinite-horizon approximations to capture long-horizon objectives. Experiments on offline GCRL benchmarks show that SVL combined with hierarchical actors matches or surpasses strong hierarchical TD and Monte Carlo baselines, excelling on complex, long-horizon tasks. Webpage and Code: https://simple-robotics.github.io/publications/survival-value-learning/
format Preprint
id arxiv_https___arxiv_org_abs_2604_17551
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SVL: Goal-Conditioned Reinforcement Learning as Survival Learning
Tiofack, Franki Nguimatsia
Schramm, Fabian
Hellard, Théotime Le
Carpentier, Justin
Machine Learning
Artificial Intelligence
Standard approaches to goal-conditioned reinforcement learning (GCRL) that rely on temporal-difference learning can be unstable and sample-inefficient due to bootstrapping. While recent work has explored contrastive and supervised formulations to improve stability, we present a probabilistic alternative, called survival value learning (SVL), that reframes GCRL as a survival learning problem by modeling the time-to-goal from each state as a probability distribution. This structured distributional Monte Carlo perspective yields a closed-form identity that expresses the goal-conditioned value function as a discounted sum of survival probabilities, enabling value estimation via a hazard model trained via maximum likelihood on both event and right-censored trajectories. We introduce three practical value estimators, including finite-horizon truncation and two binned infinite-horizon approximations to capture long-horizon objectives. Experiments on offline GCRL benchmarks show that SVL combined with hierarchical actors matches or surpasses strong hierarchical TD and Monte Carlo baselines, excelling on complex, long-horizon tasks. Webpage and Code: https://simple-robotics.github.io/publications/survival-value-learning/
title SVL: Goal-Conditioned Reinforcement Learning as Survival Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.17551