Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choi, Jinwoo, Lee, Sang-Hyun, Seo, Seung-Woo
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908808223129600
author Choi, Jinwoo
Lee, Sang-Hyun
Seo, Seung-Woo
author_facet Choi, Jinwoo
Lee, Sang-Hyun
Seo, Seung-Woo
contents Offline goal-conditioned reinforcement learning remains challenging for long-horizon tasks. While hierarchical approaches mitigate this issue by decomposing tasks, most existing methods rely on separate high- and low-level networks and generate only a single intermediate subgoal, making them inadequate for complex tasks that require coordinating multiple intermediate decisions. To address this limitation, we draw inspiration from the chain-of-thought paradigm and propose the Chain-of-Goals Hierarchical Policy (CoGHP), a novel framework that reformulates hierarchical decision-making as autoregressive sequence modeling within a unified architecture. Given a state and a final goal, CoGHP autoregressively generates a sequence of latent subgoals followed by the primitive action, where each latent subgoal acts as a reasoning step that conditions subsequent predictions. To implement this efficiently, we pioneer the use of an MLP-Mixer backbone, which supports cross-token communication and captures structural relationships among state, goal, latent subgoals, and action. Across challenging navigation and manipulation benchmarks, CoGHP consistently outperforms strong offline baselines, demonstrating improved performance on long-horizon tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03389
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
Choi, Jinwoo
Lee, Sang-Hyun
Seo, Seung-Woo
Machine Learning
Artificial Intelligence
Offline goal-conditioned reinforcement learning remains challenging for long-horizon tasks. While hierarchical approaches mitigate this issue by decomposing tasks, most existing methods rely on separate high- and low-level networks and generate only a single intermediate subgoal, making them inadequate for complex tasks that require coordinating multiple intermediate decisions. To address this limitation, we draw inspiration from the chain-of-thought paradigm and propose the Chain-of-Goals Hierarchical Policy (CoGHP), a novel framework that reformulates hierarchical decision-making as autoregressive sequence modeling within a unified architecture. Given a state and a final goal, CoGHP autoregressively generates a sequence of latent subgoals followed by the primitive action, where each latent subgoal acts as a reasoning step that conditions subsequent predictions. To implement this efficiently, we pioneer the use of an MLP-Mixer backbone, which supports cross-token communication and captures structural relationships among state, goal, latent subgoals, and action. Across challenging navigation and manipulation benchmarks, CoGHP consistently outperforms strong offline baselines, demonstrating improved performance on long-horizon tasks.
title Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03389