Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hori, Toshiaki, DeCastro, Jonathan, Gopinath, Deepak, Balachandran, Avinash, Rosman, Guy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918449288052736
author Hori, Toshiaki
DeCastro, Jonathan
Gopinath, Deepak
Balachandran, Avinash
Rosman, Guy
author_facet Hori, Toshiaki
DeCastro, Jonathan
Gopinath, Deepak
Balachandran, Avinash
Rosman, Guy
contents We propose a new approach for solving planning problems with a hierarchical structure, fusing reinforcement learning and MPC planning. Our formulation tightly and elegantly couples the two planning paradigms. It leverages reinforcement learning actions to inform the MPPI sampler, and adaptively aggregates MPPI samples to inform the value estimation. The resulting adaptive process leverages further MPPI exploration where value estimates are uncertain, and improves training robustness and the overall resulting policies. This results in a robust planning approach that can handle complex planning problems and easily adapts to different applications, as demonstrated over several domains, including race driving, modified Acrobot, and Lunar Lander with added obstacles. Our results in these domains show better data efficiency and overall performance in terms of both rewards and task success, with up to a 72% increase in success rate compared to existing approaches, as well as accelerated convergence (x2.1) compared to non-adaptive sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
Hori, Toshiaki
DeCastro, Jonathan
Gopinath, Deepak
Balachandran, Avinash
Rosman, Guy
Machine Learning
Artificial Intelligence
Robotics
We propose a new approach for solving planning problems with a hierarchical structure, fusing reinforcement learning and MPC planning. Our formulation tightly and elegantly couples the two planning paradigms. It leverages reinforcement learning actions to inform the MPPI sampler, and adaptively aggregates MPPI samples to inform the value estimation. The resulting adaptive process leverages further MPPI exploration where value estimates are uncertain, and improves training robustness and the overall resulting policies. This results in a robust planning approach that can handle complex planning problems and easily adapts to different applications, as demonstrated over several domains, including race driving, modified Acrobot, and Lunar Lander with added obstacles. Our results in these domains show better data efficiency and overall performance in terms of both rewards and task success, with up to a 72% increase in success rate compared to existing approaches, as well as accelerated convergence (x2.1) compared to non-adaptive sampling.
title Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2512.17091