End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zifan, Yang, Xun, Zhao, Jianzhuang, Zhou, Jiaming, Ma, Teli, Gao, Ziyao, Ajoudani, Arash, Liang, Junwei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918121295577088
author Wang, Zifan
Yang, Xun
Zhao, Jianzhuang
Zhou, Jiaming
Ma, Teli
Gao, Ziyao
Ajoudani, Arash
Liang, Junwei
author_facet Wang, Zifan
Yang, Xun
Zhao, Jianzhuang
Zhou, Jiaming
Ma, Teli
Gao, Ziyao
Ajoudani, Arash
Liang, Junwei
contents The deployment of humanoid robots in unstructured, human-centric environments requires navigation capabilities that extend beyond simple locomotion to include robust perception, provable safety, and socially aware behavior. Current reinforcement learning approaches are often limited by blind controllers that lack environmental awareness or by vision-based systems that fail to perceive complex 3D obstacles. In this work, we present an end-to-end locomotion policy that directly maps raw, spatio-temporal LiDAR point clouds to motor commands, enabling robust navigation in cluttered dynamic scenes. We formulate the control problem as a Constrained Markov Decision Process (CMDP) to formally separate safety from task objectives. Our key contribution is a novel methodology that translates the principles of Control Barrier Functions (CBFs) into costs within the CMDP, allowing a model-free Penalized Proximal Policy Optimization (P3O) to enforce safety constraints during training. Furthermore, we introduce a set of comfort-oriented rewards, grounded in human-robot interaction research, to promote motions that are smooth, predictable, and less intrusive. We demonstrate the efficacy of our framework through a successful sim-to-real transfer to a physical humanoid robot, which exhibits agile and safe navigation around both static and dynamic 3D obstacles.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy
Wang, Zifan
Yang, Xun
Zhao, Jianzhuang
Zhou, Jiaming
Ma, Teli
Gao, Ziyao
Ajoudani, Arash
Liang, Junwei
Robotics
The deployment of humanoid robots in unstructured, human-centric environments requires navigation capabilities that extend beyond simple locomotion to include robust perception, provable safety, and socially aware behavior. Current reinforcement learning approaches are often limited by blind controllers that lack environmental awareness or by vision-based systems that fail to perceive complex 3D obstacles. In this work, we present an end-to-end locomotion policy that directly maps raw, spatio-temporal LiDAR point clouds to motor commands, enabling robust navigation in cluttered dynamic scenes. We formulate the control problem as a Constrained Markov Decision Process (CMDP) to formally separate safety from task objectives. Our key contribution is a novel methodology that translates the principles of Control Barrier Functions (CBFs) into costs within the CMDP, allowing a model-free Penalized Proximal Policy Optimization (P3O) to enforce safety constraints during training. Furthermore, we introduce a set of comfort-oriented rewards, grounded in human-robot interaction research, to promote motions that are smooth, predictable, and less intrusive. We demonstrate the efficacy of our framework through a successful sim-to-real transfer to a physical humanoid robot, which exhibits agile and safe navigation around both static and dynamic 3D obstacles.
title End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy
topic Robotics
url https://arxiv.org/abs/2508.07611