Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bastankhah, Mahsa, Liu, Grace, Arumugam, Dilip, Griffiths, Thomas L., Eysenbach, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918161610178560
author Bastankhah, Mahsa
Liu, Grace
Arumugam, Dilip
Griffiths, Thomas L.
Eysenbach, Benjamin
author_facet Bastankhah, Mahsa
Liu, Grace
Arumugam, Dilip
Griffiths, Thomas L.
Eysenbach, Benjamin
contents In this work, we take a first step toward elucidating the mechanisms behind emergent exploration in unsupervised reinforcement learning. We study Single-Goal Contrastive Reinforcement Learning (SGCRL), a self-supervised algorithm capable of solving challenging long-horizon goal-reaching tasks without external rewards or curricula. We combine theoretical analysis of the algorithm's objective function with controlled experiments to understand what drives its exploration. We show that SGCRL maximizes implicit rewards shaped by its learned representations. These representations automatically modify the reward landscape to promote exploration before reaching the goal and exploitation thereafter. Our experiments also demonstrate that these exploration dynamics arise from learning low-rank representations of the state space rather than from neural network function approximation. Our improved understanding enables us to adapt SGCRL to perform safety-aware exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14129
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
Bastankhah, Mahsa
Liu, Grace
Arumugam, Dilip
Griffiths, Thomas L.
Eysenbach, Benjamin
Machine Learning
In this work, we take a first step toward elucidating the mechanisms behind emergent exploration in unsupervised reinforcement learning. We study Single-Goal Contrastive Reinforcement Learning (SGCRL), a self-supervised algorithm capable of solving challenging long-horizon goal-reaching tasks without external rewards or curricula. We combine theoretical analysis of the algorithm's objective function with controlled experiments to understand what drives its exploration. We show that SGCRL maximizes implicit rewards shaped by its learned representations. These representations automatically modify the reward landscape to promote exploration before reaching the goal and exploitation thereafter. Our experiments also demonstrate that these exploration dynamics arise from learning low-rank representations of the state space rather than from neural network function approximation. Our improved understanding enables us to adapt SGCRL to perform safety-aware exploration.
title Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
topic Machine Learning
url https://arxiv.org/abs/2510.14129