Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Zixuan, He, Xialin, Wang, Yen-Jen, Liao, Qiayuan, Ze, Yanjie, Li, Zhongyu, Sastry, S. Shankar, Wu, Jiajun, Sreenath, Koushil, Gupta, Saurabh, Peng, Xue Bin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912089249939456
author Chen, Zixuan
He, Xialin
Wang, Yen-Jen
Liao, Qiayuan
Ze, Yanjie
Li, Zhongyu
Sastry, S. Shankar
Wu, Jiajun
Sreenath, Koushil
Gupta, Saurabh
Peng, Xue Bin
author_facet Chen, Zixuan
He, Xialin
Wang, Yen-Jen
Liao, Qiayuan
Ze, Yanjie
Li, Zhongyu
Sastry, S. Shankar
Wu, Jiajun
Sreenath, Koushil
Gupta, Saurabh
Peng, Xue Bin
contents Reinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, such as low-pass filters and smoothness rewards, are often employed to develop policies with smooth behaviors. However, because these techniques are non-differentiable and usually require tedious tuning of a large set of hyperparameters, they tend to require extensive manual tuning for each robotic platform. To address this challenge and establish a general technique for enforcing smooth behaviors, we propose a simple and effective method that imposes a Lipschitz constraint on a learned policy, which we refer to as Lipschitz-Constrained Policies (LCP). We show that the Lipschitz constraint can be implemented in the form of a gradient penalty, which provides a differentiable objective that can be easily incorporated with automatic differentiation frameworks. We demonstrate that LCP effectively replaces the need for smoothing rewards or low-pass filters and can be easily integrated into training frameworks for many distinct humanoid robots. We extensively evaluate LCP in both simulation and real-world humanoid robots, producing smooth and robust locomotion controllers. All simulation and deployment code, along with complete checkpoints, is available on our project page: https://lipschitz-constrained-policy.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11825
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
Chen, Zixuan
He, Xialin
Wang, Yen-Jen
Liao, Qiayuan
Ze, Yanjie
Li, Zhongyu
Sastry, S. Shankar
Wu, Jiajun
Sreenath, Koushil
Gupta, Saurabh
Peng, Xue Bin
Robotics
Artificial Intelligence
Reinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, such as low-pass filters and smoothness rewards, are often employed to develop policies with smooth behaviors. However, because these techniques are non-differentiable and usually require tedious tuning of a large set of hyperparameters, they tend to require extensive manual tuning for each robotic platform. To address this challenge and establish a general technique for enforcing smooth behaviors, we propose a simple and effective method that imposes a Lipschitz constraint on a learned policy, which we refer to as Lipschitz-Constrained Policies (LCP). We show that the Lipschitz constraint can be implemented in the form of a gradient penalty, which provides a differentiable objective that can be easily incorporated with automatic differentiation frameworks. We demonstrate that LCP effectively replaces the need for smoothing rewards or low-pass filters and can be easily integrated into training frameworks for many distinct humanoid robots. We extensively evaluate LCP in both simulation and real-world humanoid robots, producing smooth and robust locomotion controllers. All simulation and deployment code, along with complete checkpoints, is available on our project page: https://lipschitz-constrained-policy.github.io.
title Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2410.11825