Learning Multi-Timescale Abstractions for Hierarchical Combinatorial Planning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Vivienne Huiling, Wang, Tinghuai, Pajarinen, Joni
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914574348845056
author Wang, Vivienne Huiling
Wang, Tinghuai
Pajarinen, Joni
author_facet Wang, Vivienne Huiling
Wang, Tinghuai
Pajarinen, Joni
contents The combination of exponentially large action spaces, stochastic dynamics, and long-horizon decision-making under limited resources makes Sequential Stochastic Combinatorial Optimization (SSCO) particularly challenging for reinforcement learning. Hierarchical Reinforcement Learning (HRL) offers a natural decomposition, but it places the high-level policy in a Semi-Markov Decision Process (SMDP) where actions have variable durations, making it difficult to learn a world model that is suitable for planning. We introduce a model-based hierarchical framework for sequential stochastic combinatorial decision-making that directly addresses this issue. Our method combines a latent-space tree-search planner with an SMDP-aware world model for variable-duration decisions. A multi-timescale objective structures the latent dynamics so that transition magnitudes reflect the effective temporal scales of abstract actions, enabling efficient lookahead under adaptive temporal abstraction. We further learn a subgoal-conditioned budget policy jointly with the world model to support context-aware resource allocation. Across challenging SSCO benchmarks, our method outperforms strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17058
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Multi-Timescale Abstractions for Hierarchical Combinatorial Planning
Wang, Vivienne Huiling
Wang, Tinghuai
Pajarinen, Joni
Machine Learning
The combination of exponentially large action spaces, stochastic dynamics, and long-horizon decision-making under limited resources makes Sequential Stochastic Combinatorial Optimization (SSCO) particularly challenging for reinforcement learning. Hierarchical Reinforcement Learning (HRL) offers a natural decomposition, but it places the high-level policy in a Semi-Markov Decision Process (SMDP) where actions have variable durations, making it difficult to learn a world model that is suitable for planning. We introduce a model-based hierarchical framework for sequential stochastic combinatorial decision-making that directly addresses this issue. Our method combines a latent-space tree-search planner with an SMDP-aware world model for variable-duration decisions. A multi-timescale objective structures the latent dynamics so that transition magnitudes reflect the effective temporal scales of abstract actions, enabling efficient lookahead under adaptive temporal abstraction. We further learn a subgoal-conditioned budget policy jointly with the world model to support context-aware resource allocation. Across challenging SSCO benchmarks, our method outperforms strong baselines.
title Learning Multi-Timescale Abstractions for Hierarchical Combinatorial Planning
topic Machine Learning
url https://arxiv.org/abs/2605.17058