Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Coholich, Jeremiah, Murtaza, Muhammad Ali, Hutchinson, Seth, Kira, Zsolt
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911021251166208
author Coholich, Jeremiah
Murtaza, Muhammad Ali
Hutchinson, Seth
Kira, Zsolt
author_facet Coholich, Jeremiah
Murtaza, Muhammad Ali
Hutchinson, Seth
Kira, Zsolt
contents We propose a novel hierarchical reinforcement learning framework for quadruped locomotion over challenging terrain. Our approach incorporates a two-layer hierarchy in which a high-level policy (HLP) selects optimal goals for a low-level policy (LLP). The LLP is trained using an on-policy actor-critic RL algorithm and is given footstep placements as goals. We propose an HLP that does not require any additional training or environment samples and instead operates via an online optimization process over the learned value function of the LLP. We demonstrate the benefits of this framework by comparing it with an end-to-end reinforcement learning (RL) approach. We observe improvements in its ability to achieve higher rewards with fewer collisions across an array of different terrains, including terrains more difficult than any encountered during training.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion
Coholich, Jeremiah
Murtaza, Muhammad Ali
Hutchinson, Seth
Kira, Zsolt
Robotics
Artificial Intelligence
We propose a novel hierarchical reinforcement learning framework for quadruped locomotion over challenging terrain. Our approach incorporates a two-layer hierarchy in which a high-level policy (HLP) selects optimal goals for a low-level policy (LLP). The LLP is trained using an on-policy actor-critic RL algorithm and is given footstep placements as goals. We propose an HLP that does not require any additional training or environment samples and instead operates via an online optimization process over the learned value function of the LLP. We demonstrate the benefits of this framework by comparing it with an end-to-end reinforcement learning (RL) approach. We observe improvements in its ability to achieve higher rewards with fewer collisions across an array of different terrains, including terrains more difficult than any encountered during training.
title Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2506.20036