Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912862339858432 |
|---|---|
| author | Kobanda, Anthony Radji, Waris Petitbois, Mathieu Maillard, Odalric-Ambrym Portelas, Rémy |
| author_facet | Kobanda, Anthony Radji, Waris Petitbois, Mathieu Maillard, Odalric-Ambrym Portelas, Rémy |
| contents | Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks remains challenging, notably due to compounding value-estimation errors. Principled geometric offers a potential solution to address these issues. Following this insight, we introduce Projective Quasimetric Planning (ProQ), a compositional framework that learns an asymmetric distance and then repurposes it, firstly as a repulsive energy forcing a sparse set of keypoints to uniformly spread over the learned latent space, and secondly as a structured directional cost guiding towards proximal sub-goals. In particular, ProQ couples this geometry with a Lagrangian out-of-distribution detector to ensure the learned keypoints stay within reachable areas. By unifying metric learning, keypoint coverage, and goal-conditioned control, our approach produces meaningful sub-goals and robustly drives long-horizon goal-reaching on diverse a navigation benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18847 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning Kobanda, Anthony Radji, Waris Petitbois, Mathieu Maillard, Odalric-Ambrym Portelas, Rémy Machine Learning Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks remains challenging, notably due to compounding value-estimation errors. Principled geometric offers a potential solution to address these issues. Following this insight, we introduce Projective Quasimetric Planning (ProQ), a compositional framework that learns an asymmetric distance and then repurposes it, firstly as a repulsive energy forcing a sparse set of keypoints to uniformly spread over the learned latent space, and secondly as a structured directional cost guiding towards proximal sub-goals. In particular, ProQ couples this geometry with a Lagrangian out-of-distribution detector to ensure the learned keypoints stay within reachable areas. By unifying metric learning, keypoint coverage, and goal-conditioned control, our approach produces meaningful sub-goals and robustly drives long-horizon goal-reaching on diverse a navigation benchmarks. |
| title | Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2506.18847 |