Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lawrence, Nathan P., Mesbah, Ali
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2512.06471
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913128050065408
author Lawrence, Nathan P.
Mesbah, Ali
author_facet Lawrence, Nathan P.
Mesbah, Ali
contents Goal-conditioned reinforcement learning (RL) concerns the problem of training an agent to maximize the probability of reaching target goal states. This paper presents an analysis of the goal-conditioned setting based on optimal control. In particular, we derive an optimality gap between more classical, often quadratic, objectives and the goal-conditioned reward, elucidating the success of goal-conditioned RL and why classical ``dense'' rewards can falter. We then consider the partially observed Markov decision setting and connect state estimation to our probabilistic reward, making the goal-conditioned reward well suited to dual control problems. The advantages of goal-conditioned policies are validated on nonlinear and uncertain environments using both RL and predictive control techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
Lawrence, Nathan P.
Mesbah, Ali
Machine Learning
Artificial Intelligence
Goal-conditioned reinforcement learning (RL) concerns the problem of training an agent to maximize the probability of reaching target goal states. This paper presents an analysis of the goal-conditioned setting based on optimal control. In particular, we derive an optimality gap between more classical, often quadratic, objectives and the goal-conditioned reward, elucidating the success of goal-conditioned RL and why classical ``dense'' rewards can falter. We then consider the partially observed Markov decision setting and connect state estimation to our probabilistic reward, making the goal-conditioned reward well suited to dual control problems. The advantages of goal-conditioned policies are validated on nonlinear and uncertain environments using both RL and predictive control techniques.
title Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.06471