Gait-Conditioned Reinforcement Learning with Multi-Phase Curriculum for Humanoid Locomotion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Tianhu, Bao, Lingfan, Zhou, Chengxu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914037515681792
author Peng, Tianhu
Bao, Lingfan
Zhou, Chengxu
author_facet Peng, Tianhu
Bao, Lingfan
Zhou, Chengxu
contents We present a unified gait-conditioned reinforcement learning framework that enables humanoid robots to perform standing, walking, running, and smooth transitions within a single recurrent policy. A compact reward routing mechanism dynamically activates gait-specific objectives based on a one-hot gait ID, mitigating reward interference and supporting stable multi-gait learning. Human-inspired reward terms promote biomechanically natural motions, such as straight-knee stance and coordinated arm-leg swing, without requiring motion capture data. A structured curriculum progressively introduces gait complexity and expands command space over multiple phases. In simulation, the policy successfully achieves robust standing, walking, running, and gait transitions. On the real Unitree G1 humanoid, we validate standing, walking, and walk-to-stand transitions, demonstrating stable and coordinated locomotion. This work provides a scalable, reference-free solution toward versatile and naturalistic humanoid control across diverse modes and environments.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gait-Conditioned Reinforcement Learning with Multi-Phase Curriculum for Humanoid Locomotion
Peng, Tianhu
Bao, Lingfan
Zhou, Chengxu
Robotics
We present a unified gait-conditioned reinforcement learning framework that enables humanoid robots to perform standing, walking, running, and smooth transitions within a single recurrent policy. A compact reward routing mechanism dynamically activates gait-specific objectives based on a one-hot gait ID, mitigating reward interference and supporting stable multi-gait learning. Human-inspired reward terms promote biomechanically natural motions, such as straight-knee stance and coordinated arm-leg swing, without requiring motion capture data. A structured curriculum progressively introduces gait complexity and expands command space over multiple phases. In simulation, the policy successfully achieves robust standing, walking, running, and gait transitions. On the real Unitree G1 humanoid, we validate standing, walking, and walk-to-stand transitions, demonstrating stable and coordinated locomotion. This work provides a scalable, reference-free solution toward versatile and naturalistic humanoid control across diverse modes and environments.
title Gait-Conditioned Reinforcement Learning with Multi-Phase Curriculum for Humanoid Locomotion
topic Robotics
url https://arxiv.org/abs/2505.20619